Case study · Reliability Sprint
From two-day deploys to fifteen minutes at a 300-engineer fintech.
How a six-week engagement rebuilt a fear-driven release process into a boring, continuous pipeline for a 300-engineer org.
The challenge
Every release was an event — and everyone dreaded it.
The team had grown from 40 to 300 engineers in two years, but the release process hadn’t changed. Deploys were batched, manual, and scheduled for Friday afternoons behind a war room. On-call was quietly burning out the senior engineers who understood the pipeline — the same people the company could least afford to lose.
We weren’t slow because our engineers were slow. We were slow because shipping was terrifying.
— VP Engineering
What we did
Made the safe path the fast path.
- Decoupled deploy from release. Feature flags let code ship continuously while releases stayed a business decision — killing the Friday war-room outright.
- Rebuilt CI for speed and trust. Cut the pipeline from 47 to 9 minutes, quarantined flaky tests, and made green actually mean green.
- One-click canary and rollback. Progressive rollout with automatic rollback on SLO breach, so any engineer could ship without a safety net of humans.
- Fixed the alerts, not just the pager. Deleted the noise, tied every remaining alert to an SLO, and handed on-call a runbook that works at 3am.
The result
Shipping got boring. The roadmap got faster.
Six weeks in, deploys were a non-event — 140+ a week, any engineer, any hour. Lead time dropped from two days to fifteen minutes and change failures fell by three-quarters. On-call load dropped enough that engineers volunteered to cover shifts again.
Most importantly, the senior engineers who’d been babysitting the pipeline went back to building product — and stayed.