Case study · Reliability Sprint

Cutting on-call load without adding headcount.

Deleted alert noise, tied every page to an SLO, and gave senior engineers their nights back.

Series-D SaaSReliability Sprint

Pages / on-call week

27 6

Alerts w/o an SLO

62 0

Median time to ack

9 min 2 min

Senior on-call regrets

high low

The challenge

On-call had become a punishment shift, not a responsibility.

Twenty-seven pages a week sounds survivable until you notice most of them fired on symptoms nobody could act on at 3am — a disk nudging a threshold, a retry storm that resolved itself in ninety seconds. Engineers stopped trusting the pager, which meant the pages that mattered got the same shrug as the noise.

What we did

Rebuilt the alert set around what a human can actually do about it.

  1. Audited every alert against a real SLO. If a page didn’t map to something a customer would notice, it got deleted or downgraded to a dashboard.
  2. Rewrote the runbooks that survived. Every remaining alert now links to a runbook with the actual fix, not a wiki page from two reorgs ago.
  3. Gave on-call authority to silence noise on sight. No approval process to mute a bad alert — if it pages and doesn’t help, it’s gone by morning.
  4. Rebalanced the rotation. Fewer, more meaningful pages meant the rotation could shrink without anyone carrying more risk.

The result

Pages per on-call week dropped from 27 to 6 — and the six that remain get taken seriously, because they’re real. Median acknowledgment time fell from nine minutes to two, simply because engineers stopped assuming it was noise.

No new hires. Just a pager worth believing again.