WM Blog · Clara

Incident Ownership Collapses at First Handoff

Engineering teams still assign incidents to the person who happens to be online rather than the function that must fix the root cause. Production keeps swinging between quick patches and repeated outages.

Fractured alert chain across deserted desks in a dark engineering office

Your engineering manager signs off on the latest on-call roster for a payments service. The document lists primary and secondary contacts for every hour, yet no one owns the outcome once the alert routes past the first responder.

A Friday 4:30pm database timeout lands with the junior SRE on shift. She pages the database lead, who has already left for the airport. The ticket sits open while revenue stalls.

The next morning the CFO sees the outage report and asks who approved the rollback. Three teams claim the change originated elsewhere. No single owner can be identified because the escalation matrix stops at the initial assignee.

This pattern repeats across mid-sized Australian SaaS firms that copied enterprise runbooks without rewriting decision rights. The result is faster detection but slower resolution every time the incident crosses a team boundary.

Procurement and finance feel the impact first when failed transactions trigger manual refunds and support tickets. The engineering group reports 99.7 percent uptime while cash collection drops for the quarter.

Fix the model by tying every production alert to a named function with budget and authority to change code or contracts. Stop measuring response time alone and start measuring time to restored revenue.

Without that shift, Friday escalations will keep landing on whoever is cheapest to disturb rather than whoever can actually close the loop.

Operating Models Escalation Incident Response Engineering Leadership