Incident Response Runbook
Execute structured incident response for: {{ incident_title }}
Severity: {{ severity }} | Service: {{ affected_service }}
Severity Decision Matrix
| Signal |
SEV1 |
SEV2 |
SEV3 |
SEV4 |
| User impact |
>25% users |
5-25% users |
<5% users |
None |
| Revenue impact |
High / checkout broken |
Medium / degraded |
Low / workaround exists |
None |
| Data integrity |
At risk |
Not at risk |
Not at risk |
Not at risk |
| Response |
All-hands, immediate |
Team, within 15 min |
Individual, within 1h |
Business hours |
| Post-mortem |
Required within 48h |
Required within 48h |
Brief review within 1 week |
Ticket only |
Severity-specific playbooks with escalation matrices
Workflow
Phase 1 — DETECT & DECLARE (0-5 min)
Confirm incident is real — not a monitoring false positive
- Check dashboard for corroborating signals
- Verify user impact exists (not just synthetic test failure)
Declare the incident
- Assign Incident Commander (IC) — single decision-maker
- Open incident channel:
{{ incident_channel }}
- Post initial broadcast:
INCIDENT DECLARED — {{ severity }}
Title: {{ incident_title }}
Affected: {{ affected_service }}
IC: [your name]
Bridge: {{ incident_channel }}
Status: INVESTIGATING
Page on-call if not already paged
- Escalate to secondary if primary does not ACK within 5 min (SEV1) / 15 min (SEV2)
Phase 2 — TRIAGE (5-15 min)
Answer ALL five triage questions — do not skip any:
| # |
Question |
What to check |
| 1 |
WHAT is broken? |
Services, endpoints, error type (5xx/timeout/data loss) |
| 2 |
WHO is impacted? |
% users, segments, regions, internal vs customer-facing |
| 3 |
HOW LONG? |
Start time, ongoing vs intermittent |
| 4 |
WHAT CHANGED? |
Deployments (last 2h), infra changes, traffic patterns, third-party status |
| 5 |
CAN WE MITIGATE NOW? |
Kill switch, feature flag, rollback, traffic reroute |
Phase 3 — INVESTIGATE (15-60 min)
Systematic investigation order:
- Dashboards — error rate, latency, throughput
- Logs — first error occurrence, error patterns
- Recent changes — deployment history, config changes
- Dependencies — database, cache, external API status
- Capacity — CPU, memory, disk, connection pool
Investigation log (IC maintains):
[HH:MM] SIGNAL: [what was observed]
[HH:MM] ACTION: [what was tried]
[HH:MM] RESULT: [what happened]
[HH:MM] HYPOTHESIS: [current best guess]
[HH:MM] REJECTED: [why hypothesis was wrong] ← capture failed hypotheses too
Status updates (every 15 min to {{ incident_channel }}):
STATUS UPDATE — {{ severity }}
Elapsed: [X] minutes
Status: INVESTIGATING / MITIGATING / MONITORING
Hypothesis: [one sentence]
Next action: [what is being tried]
ETA: [estimate or "unknown"]
Phase 4 — MITIGATE (variable)
Apply mitigation in risk-priority order (lowest risk first):
| Priority |
Action |
Risk |
Reversibility |
| 1 |
Kill switch / feature flag |
Lowest |
Instant |
| 2 |
Traffic routing to healthy region |
Low |
Fast |
| 3 |
Rollback deployment |
Medium |
Minutes |
| 4 |
Scale up capacity |
Medium |
Minutes |
| 5 |
Infrastructure fix (restart, clear cache) |
Medium |
Variable |
| 6 |
Hotfix deployment |
Highest |
Slow |
Document the decision:
MITIGATION: [action]
Reason: [why chosen over alternatives]
Risk: [what could go wrong]
Rollback plan: [how to undo if it makes things worse]
Approved by IC at [HH:MM]
Phase 5 — RESOLVE
Incident is resolved when ALL conditions met:
- Error rate back to baseline (<1% or pre-incident level)
- Latency back to baseline
- All services reporting healthy
- No new customer complaints
Resolution announcement:
INCIDENT RESOLVED — {{ severity }}
Title: {{ incident_title }}
Duration: [X hours Y minutes]
Root Cause: [one sentence]
Resolution: [what fixed it]
Post-mortem: Scheduled for [date]
Post-resolution checklist:
Phase 6 — POST-MORTEM
Create within 48h for SEV1/SEV2. Structure:
INCIDENT POST-MORTEM
Title: {{ incident_title }}
Date: [date] | Duration: [total] | Severity: {{ severity }}
IMPACT: [users affected, features degraded, revenue impact]
TIMELINE: [chronological key events]
ROOT CAUSE: [technical cause — be specific, avoid blame]
CONTRIBUTING FACTORS: [systemic issues]
WHAT WENT WELL: [things that worked]
WHAT TO IMPROVE: [things that slowed response]
ACTION ITEMS:
| Action | Owner | Due Date | Priority |
|--------|-------|----------|----------|
Counter-Rationalizations
| Shortcut |
Counter |
Why |
| "Skip triage, we know what's wrong" |
Complete all 5 triage questions |
Triage reveals blast radius — you can't mitigate what you haven't measured |
| "Just restart the service" |
Investigate before restarting |
Restart masks root cause, may cause data loss, and delays actual fix |
| "We don't need an IC" |
Always assign an IC |
Without single decision-maker, conflicting actions extend the incident |
| "Post-mortem can wait" |
Schedule within 48h |
Details fade quickly; action items lose urgency after a week |
| "This is only SEV3, skip the process" |
Adapt the process, don't skip it |
SEV3s become SEV1s when unmanaged — the process scales down, not off |
| "I'll check metrics later" |
Check dashboards first in investigation |
Silent failures are invisible without metrics; logs alone miss the big picture |
| "The fix is obvious, skip investigation" |
Document hypothesis before applying fix |
"Obvious" fixes that are wrong make the incident worse and longer |
Output Format
Produce a real-time incident log:
- Header — title, severity, IC, start time
- Triage answers — all 5 questions answered
- Investigation findings — signals, hypotheses, evidence
- Mitigation decision — action, reasoning, rollback plan
- Resolution — what fixed it, duration, next steps
- Post-mortem draft — on request
References
1---2name: incident-response-runbook3description: Use when a production incident is declared, when investigating service degradation, or when validating incident response readiness. Structured workflow covering detection, triage, investigation, mitigation, resolution, and post-mortem. Integrates with PagerDuty, Slack, Jira, and Confluence for coordinated response.4---56# Incident Response Runbook78Execute structured incident response for: **{{ incident_title }}**9Severity: **{{ severity }}** | Service: **{{ affected_service }}**1011## Severity Decision Matrix1213| Signal | SEV1 | SEV2 | SEV3 | SEV4 |14|--------|------|------|------|------|15| User impact | >25% users | 5-25% users | <5% users | None |16| Revenue impact | High / checkout broken | Medium / degraded | Low / workaround exists | None |17| Data integrity | At risk | Not at risk | Not at risk | Not at risk |18| Response | All-hands, immediate | Team, within 15 min | Individual, within 1h | Business hours |19| Post-mortem | Required within 48h | Required within 48h | Brief review within 1 week | Ticket only |2021[Severity-specific playbooks with escalation matrices](./references/severity-playbooks.md)2223## Workflow2425### Phase 1 — DETECT & DECLARE (0-5 min)26271. **Confirm incident is real** — not a monitoring false positive28 - Check dashboard for corroborating signals29 - Verify user impact exists (not just synthetic test failure)30312. **Declare the incident**32 - Assign Incident Commander (IC) — single decision-maker33 - Open incident channel: `{{ incident_channel }}`34 - Post initial broadcast:3536 ```37 INCIDENT DECLARED — {{ severity }}38 Title: {{ incident_title }}39 Affected: {{ affected_service }}40 IC: [your name]41 Bridge: {{ incident_channel }}42 Status: INVESTIGATING43 ```44453. **Page on-call** if not already paged46 - Escalate to secondary if primary does not ACK within 5 min (SEV1) / 15 min (SEV2)4748### Phase 2 — TRIAGE (5-15 min)4950Answer ALL five triage questions — do not skip any:5152| # | Question | What to check |53|---|----------|--------------|54| 1 | **WHAT** is broken? | Services, endpoints, error type (5xx/timeout/data loss) |55| 2 | **WHO** is impacted? | % users, segments, regions, internal vs customer-facing |56| 3 | **HOW LONG?** | Start time, ongoing vs intermittent |57| 4 | **WHAT CHANGED?** | Deployments (last 2h), infra changes, traffic patterns, third-party status |58| 5 | **CAN WE MITIGATE NOW?** | Kill switch, feature flag, rollback, traffic reroute |5960### Phase 3 — INVESTIGATE (15-60 min)6162Systematic investigation order:631. **Dashboards** — error rate, latency, throughput642. **Logs** — first error occurrence, error patterns653. **Recent changes** — deployment history, config changes664. **Dependencies** — database, cache, external API status675. **Capacity** — CPU, memory, disk, connection pool6869**Investigation log** (IC maintains):70```71[HH:MM] SIGNAL: [what was observed]72[HH:MM] ACTION: [what was tried]73[HH:MM] RESULT: [what happened]74[HH:MM] HYPOTHESIS: [current best guess]75[HH:MM] REJECTED: [why hypothesis was wrong] ← capture failed hypotheses too76```7778**Status updates** (every 15 min to `{{ incident_channel }}`):79```80STATUS UPDATE — {{ severity }}81Elapsed: [X] minutes82Status: INVESTIGATING / MITIGATING / MONITORING83Hypothesis: [one sentence]84Next action: [what is being tried]85ETA: [estimate or "unknown"]86```8788### Phase 4 — MITIGATE (variable)8990Apply mitigation in **risk-priority order** (lowest risk first):9192| Priority | Action | Risk | Reversibility |93|----------|--------|------|---------------|94| 1 | Kill switch / feature flag | Lowest | Instant |95| 2 | Traffic routing to healthy region | Low | Fast |96| 3 | Rollback deployment | Medium | Minutes |97| 4 | Scale up capacity | Medium | Minutes |98| 5 | Infrastructure fix (restart, clear cache) | Medium | Variable |99| 6 | Hotfix deployment | Highest | Slow |100101**Document the decision:**102```103MITIGATION: [action]104Reason: [why chosen over alternatives]105Risk: [what could go wrong]106Rollback plan: [how to undo if it makes things worse]107Approved by IC at [HH:MM]108```109110### Phase 5 — RESOLVE111112**Incident is resolved when ALL conditions met:**113- Error rate back to baseline (<1% or pre-incident level)114- Latency back to baseline115- All services reporting healthy116- No new customer complaints117118**Resolution announcement:**119```120INCIDENT RESOLVED — {{ severity }}121Title: {{ incident_title }}122Duration: [X hours Y minutes]123Root Cause: [one sentence]124Resolution: [what fixed it]125Post-mortem: Scheduled for [date]126```127128**Post-resolution checklist:**129- [ ] Update status page130- [ ] Notify customer success team131- [ ] Create post-mortem ticket (Jira/Linear)132- [ ] Schedule post-mortem meeting (within 48h for SEV1/SEV2)133- [ ] Close PagerDuty incident134135### Phase 6 — POST-MORTEM136137Create within 48h for SEV1/SEV2. Structure:138139```140INCIDENT POST-MORTEM141Title: {{ incident_title }}142Date: [date] | Duration: [total] | Severity: {{ severity }}143144IMPACT: [users affected, features degraded, revenue impact]145TIMELINE: [chronological key events]146ROOT CAUSE: [technical cause — be specific, avoid blame]147CONTRIBUTING FACTORS: [systemic issues]148WHAT WENT WELL: [things that worked]149WHAT TO IMPROVE: [things that slowed response]150151ACTION ITEMS:152| Action | Owner | Due Date | Priority |153|--------|-------|----------|----------|154```155156## Counter-Rationalizations157158| Shortcut | Counter | Why |159|----------|---------|-----|160| "Skip triage, we know what's wrong" | Complete all 5 triage questions | Triage reveals blast radius — you can't mitigate what you haven't measured |161| "Just restart the service" | Investigate before restarting | Restart masks root cause, may cause data loss, and delays actual fix |162| "We don't need an IC" | Always assign an IC | Without single decision-maker, conflicting actions extend the incident |163| "Post-mortem can wait" | Schedule within 48h | Details fade quickly; action items lose urgency after a week |164| "This is only SEV3, skip the process" | Adapt the process, don't skip it | SEV3s become SEV1s when unmanaged — the process scales down, not off |165| "I'll check metrics later" | Check dashboards first in investigation | Silent failures are invisible without metrics; logs alone miss the big picture |166| "The fix is obvious, skip investigation" | Document hypothesis before applying fix | "Obvious" fixes that are wrong make the incident worse and longer |167168## Output Format169170Produce a real-time incident log:1711. **Header** — title, severity, IC, start time1722. **Triage answers** — all 5 questions answered1733. **Investigation findings** — signals, hypotheses, evidence1744. **Mitigation decision** — action, reasoning, rollback plan1755. **Resolution** — what fixed it, duration, next steps1766. **Post-mortem draft** — on request177178## References179180- [Severity Playbooks — Escalation Matrices & Investigation Runbooks](./references/severity-playbooks.md)