For engineers who own production: run incidents calmly and defend reliability on-call.
Reach for this when you own a service on-call and want incidents to be routine, not chaos. It carries you through the whole lifecycle - instrument and de-noise your signals, call severity fast, work the runbook, keep customers honestly informed on the status page and run the external crisis comms (first statement inside the hour, speculation never), hand the shift off with a completeness checklist, write the blameless postmortem on the 48-hour clock with the five-whys-stops-at-process rule, and turn SLO error budgets into automatic team decisions. Templates and severity thresholds included at every step. Built for engineers and SREs who carry the pager.
Click to play with sound.
Arranged in the author's recommended order. Walk through them in sequence, or open any one on its own.
Writes an operational runbook for a service covering symptoms, diagnostic checks, mitigations, and escalation paths. Use when shipping a new service or when an existing service lacks incident procedures.
View skill