Skip to content
status: steady

Illustrative composite · not a client claim

On-call that pages all night

Twenty-five pages per engineer per week, most acknowledged and ignored. Real incidents hide inside the noise. Postmortems produce action items nobody has time for, and the people carrying pagers have started interviewing elsewhere.

What we investigate

  • Every alert that fired in the last ninety days, and what happened next
  • Which alerts map to symptoms users feel, and which are internal noise
  • Which dashboards exist and which anyone reads
  • The real reliability requirement per service, not the assumed one
  • The incident process from first page to postmortem

Typical deliverables

  • An alert inventory with a keep, downgrade, or delete decision per rule
  • SLOs for the user journeys that matter, wired to burn-rate alerts
  • A runbook for every page that survives the cull
  • A calmer rotation, measured before and after

What typically gets better

  • A page comes to mean something again
  • The rotation stops being a reason people interview elsewhere
  • Postmortem actions land, because there are few enough to take seriously

This is Cloud & Infrastructure work. The engagement model is on How we work.

Request a call

Living this one right now?

A short email about your setup and what's not working is the best start. We'll say plainly whether we can help, and what it would take.

hello@farzanfa.com · +91 88918 87223 · Kerala, India · working worldwide · reply within one business day