Postmortem writing
Principles
- Blameless — focus on systems and process, not individuals.
- Accurate — timeline in UTC with evidence (logs, deploys, tickets).
- Actionable — every finding maps to owner + due date or ticket.
Required sections
- Summary — what happened, duration, user/business impact
- Timeline — detection, escalation, mitigation, recovery
- Root cause — technical chain; distinguish trigger vs underlying gap
- Contributing factors — process, monitoring, dependency, change management
- What went well
- Action items — prevent recurrence, improve detection, reduce blast radius
Timeline table
| Time (UTC) | Event |
|---|---|
| ... | Alert fired |
| ... | Mitigation applied |
| ... | Service restored |
Action item quality
| Weak | Strong |
|---|---|
| "Be more careful" | "Add alert on queue depth > X with runbook" |
| "Improve testing" | "Integration test for failover path in service Y" |
Severity-appropriate depth
- SEV1/2: full postmortem + review meeting.
- SEV3/4: lightweight doc still with timeline and actions.
Template
See reference.md for full markdown template.
Follow-up
Track action items to closure; link superseding ADRs or runbook updates.