Incident Response
Incidents are inevitable in any production system. The difference between a minor blip and a catastrophic failure is how prepared your team is to detect, respond, communicate, and learn. This guide covers the full incident lifecycle from classification through postmortem, with concrete templates and integration patterns.
Severity Classification
Every incident must be classified immediately. Severity determines response speed, communication cadence, and escalation paths.
| Severity |
Impact |
Examples |
Response Time |
Duration Target |
| SEV1 |
Complete outage or data loss |
Service down for all users, data corruption, security breach |
5 min |
Mitigate < 1 hour |
| SEV2 |
Major degradation |
Core feature broken for >10% users, payment failures |
15 min |
Mitigate < 4 hours |
| SEV3 |
Minor degradation |
Non-critical feature broken, elevated error rate |
1 hour |
Resolve < 24 hours |
| SEV4 |
Cosmetic or low impact |
UI glitch, misleading error message |
Next business day |
Resolve < 1 week |
Classification Decision Tree
Is the service completely unavailable to all users? -> SEV1
Is there a security breach or data exposure? -> SEV1
Is there data loss or corruption? -> SEV1
Is a core revenue feature broken for >10% users? -> SEV2
Is there financial impact (failed payments)? -> SEV2
Is a core feature broken for <10% users? -> SEV3
Is it a non-critical feature degradation? -> SEV3
Everything else -> SEV4
On-Call Setup
| Component |
Recommendation |
Rationale |
| Rotation size |
5-8 engineers minimum |
Allows 1-week shifts with recovery time |
| Shift length |
1 week (Mon 09:00 to Mon 09:00) |
Predictable, long enough for context |
| Primary + Secondary |
Always two people on call |
Backup for escalation or unavailability |
| Handoff meeting |
30 min at rotation start |
Review active issues, recent changes, known risks |
| Follow-the-sun |
Split by timezone for global teams |
No one wakes up at 3 AM regularly |
| Compensation |
On-call stipend + incident bonus |
Recognizes the burden fairly |
PagerDuty Schedule and Escalation (Terraform)
resource "pagerduty_schedule" "primary_oncall" {
name = "Platform Primary On-Call"
time_zone = "America/New_York"
layer {
name = "Primary"
start = "2025-01-06T09:00:00-05:00"
rotation_virtual_start = "2025-01-06T09:00:00-05:00"
rotation_turn_length_seconds = 604800 # 1 week
users = [
pagerduty_user.alice.id, pagerduty_user.bob.id,
pagerduty_user.carol.id, pagerduty_user.dave.id,
pagerduty_user.eve.id,
]
}
}
resource "pagerduty_escalation_policy" "platform" {
name = "Platform Escalation"
num_loops = 2
rule {
escalation_delay_in_minutes = 5
target { type = "schedule_reference"; id = pagerduty_schedule.primary_oncall.id }
}
rule {
escalation_delay_in_minutes = 10
target { type = "schedule_reference"; id = pagerduty_schedule.secondary_oncall.id }
}
rule {
escalation_delay_in_minutes = 15
target { type = "user_reference"; id = pagerduty_user.engineering_manager.id }
}
}
War Room Protocol
When a SEV1 or SEV2 is declared, open a war room -- a structured environment for incident resolution.
War Room Flow
1. DETECT Alert fires -> On-call acknowledges within 5 min
2. TRIAGE Classify severity. SEV1/SEV2 -> open war room
3. ASSEMBLE IC assigned. Slack: #inc-YYYYMMDD-description. Video bridge opened.
4. ROLES IC | Comms Lead | Operations Lead | Scribe | SMEs
5. INVESTIGATE What changed? What is blast radius? What do signals show?
6. MITIGATE Priority: restore service (rollback, feature flag, scale, failover)
7. RESOLVE Service stable. IC declares resolved.
8. FOLLOW-UP Postmortem within 48 hours. Action items tracked.
Incident Channel Template
**INCIDENT DECLARED**
Severity: SEV1 | Title: Order processing failing for all users
Detected: 2025-03-15 14:32 UTC
Impact: All users unable to complete checkout
**ROLES** IC: @alice | Comms: @bob | Ops: @carol | Scribe: @dave
**LINKS**
Status Page: https://status.example.com
Runbook: https://wiki.internal/runbooks/order-processing
Dashboard: https://grafana.internal/d/orders-overview
**TIMELINE**
14:32 - Alert fired: HighErrorBudgetBurnRate_Fast
14:35 - On-call acknowledged, war room opened
14:38 - Identified: deploy changed payment gateway config
14:42 - Rollback initiated
14:47 - Rollback complete, error rate dropping
14:55 - Fully restored | 15:00 - Resolved
Runbook Template
Every alert should link to a runbook with diagnosis, mitigation, and recovery steps.
# runbooks/order-service-high-error-rate.yml
metadata:
title: "Order Service High Error Rate"
service: order-service
severity: SEV1/SEV2
owner: platform-team
alert_names: [HighErrorBudgetBurnRate_Fast, HighErrorBudgetBurnRate_Slow]
impact: "Users cannot checkout. Revenue impact ~$2,400/min at peak."
diagnosis:
- step: Check recent deployments
command: "kubectl -n production rollout history deployment/order-service"
expected: "If deploy correlates with error onset, proceed to rollback"
- step: Check error logs
command: '{service="order-service"} |= "error" | json | level="error"'
expected: "Identify error type: database, upstream, or application"
- step: Check downstream dependencies
command: "curl -s https://payment-service.internal/healthz | jq ."
expected: "All healthy. If not, see payment-service-down runbook"
- step: Check database performance
command: "psql -h orders-db -c \"SELECT pid, state, query FROM pg_stat_activity WHERE state != 'idle' ORDER BY query_start LIMIT 20;\""
expected: "No long-running queries or lock contention"
mitigation:
- option: Rollback last deployment
when: "Error onset correlates with a deployment"
command: "kubectl -n production rollout undo deployment/order-service"
verify: "Error rate returns to baseline within 5 minutes"
- option: Scale up
when: "Capacity-related errors (connection pool, CPU)"
command: "kubectl -n production scale deployment/order-service --replicas=10"
verify: "Connection pool usage drops below 80%"
- option: Circuit breaker
when: "Payment service is root cause"
command: "kubectl set env deployment/order-service PAYMENT_CIRCUIT_BREAKER=open"
verify: "5xx rate drops, orders queue for retry"
- option: Database failover
when: "Primary database unresponsive"
command: "aws rds failover-db-cluster --db-cluster-identifier orders-cluster"
verify: "Connections re-establish within 60 seconds"
recovery:
- "Confirm baseline error rate for 15 minutes"
- "Check for data inconsistencies from failed transactions"
- "Update status page to resolved"
- "Schedule postmortem within 48 hours"
Escalation Policies
| Trigger |
Action |
Timeout |
| Alert fires |
Page primary on-call |
-- |
| No ack in 5 min |
Escalate to secondary |
5 min |
| No ack in 15 min |
Escalate to engineering manager |
10 min |
| SEV1 declared |
Auto-notify VP Eng + CTO |
Immediate |
| 30 min without mitigation |
IC requests additional responders |
IC decision |
| Customer data exposed |
Notify Security + Legal |
Immediate |
Communication Templates
# Status Page -- Investigating
**[Investigating] Elevated error rates on checkout**
We are investigating errors during checkout. Our team is engaged.
Update within 30 minutes.
# Status Page -- Identified
**[Identified] Checkout errors caused by payment config issue**
Root cause identified. Fix deploying now. Update in 15 minutes.
# Status Page -- Resolved
**[Resolved] Checkout errors resolved**
Configuration rolled back at 14:47 UTC. All systems normal.
Failed orders auto-retried. Full report within 48 hours.
# Internal Update (Slack #incidents)
**SEV1 Update -- 14:45 UTC**
Impact: Checkout down since 14:25 | Root cause: bad config deploy
Action: Rollback in progress, ETA 5 min
Revenue impact: ~$12,000 est | IC: @alice
PagerDuty / OpsGenie Integration
PagerDuty Event API v2
async function triggerIncident(params: {
title: string;
severity: 'critical' | 'error' | 'warning' | 'info';
service: string;
dedupKey: string;
}): Promise<string> {
const response = await fetch('https://events.pagerduty.com/v2/enqueue', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
routing_key: process.env.PAGERDUTY_ROUTING_KEY,
event_action: 'trigger',
dedup_key: params.dedupKey,
payload: {
summary: params.title,
severity: params.severity,
source: params.service,
},
links: [
{ href: `https://grafana.internal/d/${params.service}`, text: 'Dashboard' },
{ href: `https://wiki.internal/runbooks/${params.service}`, text: 'Runbook' },
],
}),
});
return (await response.json()).dedup_key;
}
async function resolveIncident(dedupKey: string): Promise<void> {
await fetch('https://events.pagerduty.com/v2/enqueue', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
routing_key: process.env.PAGERDUTY_ROUTING_KEY,
event_action: 'resolve',
dedup_key: dedupKey,
}),
});
}
Postmortem Template
# Postmortem: [Incident Title]
**Date:** YYYY-MM-DD | **Duration:** HH:MM | **Severity:** SEVN
**IC:** [Name] | **Authors:** [Names] | **Status:** Draft / Complete
## Summary
One paragraph: what happened, impact, resolution.
## Impact
| Metric | Value |
|--------|-------|
| User impact duration | X hours Y minutes |
| Users affected | N (Z% of total) |
| Revenue impact | $X,XXX estimated |
| SLA impact | X min against 99.9% target |
## Timeline (UTC)
| Time | Event |
|------|-------|
| 14:25 | Deploy #4521 pushed (config change) |
| 14:32 | Alert fires: HighErrorBudgetBurnRate_Fast |
| 14:36 | SEV1 declared, war room opened |
| 14:38 | Root cause: payment timeout changed 30s -> 3s |
| 14:42 | Rollback initiated |
| 14:47 | Rollback complete | 14:55 | Fully stable |
## Root Cause
[What broke and why -- focus on systems, not individuals]
## Lessons Learned
### What went well
- [Bullet points]
### What went poorly
- [Bullet points]
### Where we got lucky
- [Bullet points]
## Action Items
| ID | Action | Priority | Owner | Due | Status |
|----|--------|----------|-------|-----|--------|
| 1 | [Action] | P1 | [Name] | [Date] | Open |
Blameless Postmortem Culture
| Principle |
In Practice |
| People are not the root cause |
"The pipeline allowed unsafe config" not "Alice deployed bad config" |
| Focus on systems |
Identify process gaps, missing guardrails, tooling deficiencies |
| Assume good intentions |
Everyone tried to do the right thing with available info |
| No counterfactuals |
Not "if only X..." but "what system change prevents this?" |
| Share widely |
Postmortems are learning, not shame |
| Track to completion |
Postmortems without follow-through teach nothing |
Incident Timeline Reconstruction Sources
| Source |
What It Provides |
| Alertmanager / PagerDuty |
Alert fire/resolve times, ack times, escalations |
| Slack |
Human decisions, observations, comms |
| Git / CI |
Deploy times, code changes |
| Grafana / Metrics |
Anomaly onset, metric correlation |
| Application logs |
Error details, trace context |
| Kubernetes events |
Pod restarts, OOM kills, scheduling |
| Cloud provider |
Infrastructure changes, regional outages |
Anti-Patterns
| Anti-Pattern |
Problem |
Fix |
| No severity classification |
Every incident treated the same |
Define and enforce severity matrix |
| Hero culture |
One person handles all incidents, burns out |
Build rotation with 5+ engineers, document in runbooks |
| Blame-driven postmortems |
People hide mistakes, learning stops |
Enforce blameless process, focus on systems |
| No runbooks |
Responders waste 20+ minutes figuring out what to do |
Require runbook link on every alert |
| Postmortems without action items |
Same incident recurs |
Track items in sprint backlog with owners and deadlines |
| Alert without context |
"Check failed" with no links |
Include dashboard, runbook, impact in every alert |
| No communication plan |
Stakeholders flood war room |
Assign comms lead, use status page templates |
| Skipping postmortem for small incidents |
Miss patterns that compound |
Postmortem SEV1-2, lightweight review SEV3 |
| Testing in prod without rollback plan |
"Quick fix" makes things worse |
Always have rollback command ready first |
| Ignoring near-misses |
Only learning from actual incidents |
Track and review near-misses monthly |
| War room without clear roles |
Everyone talks, nobody acts |
Assign IC, Comms, Ops, Scribe at start |
| Over-classifying severity |
Everything is SEV1, diluted response |
Calibrate quarterly, push back on inflation |
Incident Readiness Checklist
Infrastructure
Process
Runbooks
Practice
1---2name: incident-response-173description: Provides incident response best practices covering severity classification, on-call rotation, war room protocols, runbook templates, escalation policies, and blameless postmortems. Use when handling an incident, setting up on-call, writing a postmortem, creating a runbook, configuring PagerDuty or OpsGenie, or building incident management processes.4---56# Incident Response78Incidents are inevitable in any production system. The difference between a minor blip and a catastrophic failure is how prepared your team is to detect, respond, communicate, and learn. This guide covers the full incident lifecycle from classification through postmortem, with concrete templates and integration patterns.910## Severity Classification1112Every incident must be classified immediately. Severity determines response speed, communication cadence, and escalation paths.1314| Severity | Impact | Examples | Response Time | Duration Target |15|----------|--------|----------|---------------|-----------------|16| SEV1 | Complete outage or data loss | Service down for all users, data corruption, security breach | 5 min | Mitigate < 1 hour |17| SEV2 | Major degradation | Core feature broken for >10% users, payment failures | 15 min | Mitigate < 4 hours |18| SEV3 | Minor degradation | Non-critical feature broken, elevated error rate | 1 hour | Resolve < 24 hours |19| SEV4 | Cosmetic or low impact | UI glitch, misleading error message | Next business day | Resolve < 1 week |2021### Classification Decision Tree2223```24Is the service completely unavailable to all users? -> SEV125Is there a security breach or data exposure? -> SEV126Is there data loss or corruption? -> SEV127Is a core revenue feature broken for >10% users? -> SEV228Is there financial impact (failed payments)? -> SEV229Is a core feature broken for <10% users? -> SEV330Is it a non-critical feature degradation? -> SEV331Everything else -> SEV432```3334## On-Call Setup3536| Component | Recommendation | Rationale |37|-----------|---------------|-----------|38| Rotation size | 5-8 engineers minimum | Allows 1-week shifts with recovery time |39| Shift length | 1 week (Mon 09:00 to Mon 09:00) | Predictable, long enough for context |40| Primary + Secondary | Always two people on call | Backup for escalation or unavailability |41| Handoff meeting | 30 min at rotation start | Review active issues, recent changes, known risks |42| Follow-the-sun | Split by timezone for global teams | No one wakes up at 3 AM regularly |43| Compensation | On-call stipend + incident bonus | Recognizes the burden fairly |4445### PagerDuty Schedule and Escalation (Terraform)4647```hcl48resource "pagerduty_schedule" "primary_oncall" {49 name = "Platform Primary On-Call"50 time_zone = "America/New_York"5152 layer {53 name = "Primary"54 start = "2025-01-06T09:00:00-05:00"55 rotation_virtual_start = "2025-01-06T09:00:00-05:00"56 rotation_turn_length_seconds = 604800 # 1 week57 users = [58 pagerduty_user.alice.id, pagerduty_user.bob.id,59 pagerduty_user.carol.id, pagerduty_user.dave.id,60 pagerduty_user.eve.id,61 ]62 }63}6465resource "pagerduty_escalation_policy" "platform" {66 name = "Platform Escalation"67 num_loops = 26869 rule {70 escalation_delay_in_minutes = 571 target { type = "schedule_reference"; id = pagerduty_schedule.primary_oncall.id }72 }73 rule {74 escalation_delay_in_minutes = 1075 target { type = "schedule_reference"; id = pagerduty_schedule.secondary_oncall.id }76 }77 rule {78 escalation_delay_in_minutes = 1579 target { type = "user_reference"; id = pagerduty_user.engineering_manager.id }80 }81}82```8384## War Room Protocol8586When a SEV1 or SEV2 is declared, open a war room -- a structured environment for incident resolution.8788### War Room Flow8990```911. DETECT Alert fires -> On-call acknowledges within 5 min922. TRIAGE Classify severity. SEV1/SEV2 -> open war room933. ASSEMBLE IC assigned. Slack: #inc-YYYYMMDD-description. Video bridge opened.944. ROLES IC | Comms Lead | Operations Lead | Scribe | SMEs955. INVESTIGATE What changed? What is blast radius? What do signals show?966. MITIGATE Priority: restore service (rollback, feature flag, scale, failover)977. RESOLVE Service stable. IC declares resolved.988. FOLLOW-UP Postmortem within 48 hours. Action items tracked.99```100101### Incident Channel Template102103```104**INCIDENT DECLARED**105Severity: SEV1 | Title: Order processing failing for all users106Detected: 2025-03-15 14:32 UTC107Impact: All users unable to complete checkout108109**ROLES** IC: @alice | Comms: @bob | Ops: @carol | Scribe: @dave110111**LINKS**112Status Page: https://status.example.com113Runbook: https://wiki.internal/runbooks/order-processing114Dashboard: https://grafana.internal/d/orders-overview115116**TIMELINE**11714:32 - Alert fired: HighErrorBudgetBurnRate_Fast11814:35 - On-call acknowledged, war room opened11914:38 - Identified: deploy changed payment gateway config12014:42 - Rollback initiated12114:47 - Rollback complete, error rate dropping12214:55 - Fully restored | 15:00 - Resolved123```124125## Runbook Template126127Every alert should link to a runbook with diagnosis, mitigation, and recovery steps.128129```yaml130# runbooks/order-service-high-error-rate.yml131metadata:132 title: "Order Service High Error Rate"133 service: order-service134 severity: SEV1/SEV2135 owner: platform-team136 alert_names: [HighErrorBudgetBurnRate_Fast, HighErrorBudgetBurnRate_Slow]137138impact: "Users cannot checkout. Revenue impact ~$2,400/min at peak."139140diagnosis:141 - step: Check recent deployments142 command: "kubectl -n production rollout history deployment/order-service"143 expected: "If deploy correlates with error onset, proceed to rollback"144145 - step: Check error logs146 command: '{service="order-service"} |= "error" | json | level="error"'147 expected: "Identify error type: database, upstream, or application"148149 - step: Check downstream dependencies150 command: "curl -s https://payment-service.internal/healthz | jq ."151 expected: "All healthy. If not, see payment-service-down runbook"152153 - step: Check database performance154 command: "psql -h orders-db -c \"SELECT pid, state, query FROM pg_stat_activity WHERE state != 'idle' ORDER BY query_start LIMIT 20;\""155 expected: "No long-running queries or lock contention"156157mitigation:158 - option: Rollback last deployment159 when: "Error onset correlates with a deployment"160 command: "kubectl -n production rollout undo deployment/order-service"161 verify: "Error rate returns to baseline within 5 minutes"162163 - option: Scale up164 when: "Capacity-related errors (connection pool, CPU)"165 command: "kubectl -n production scale deployment/order-service --replicas=10"166 verify: "Connection pool usage drops below 80%"167168 - option: Circuit breaker169 when: "Payment service is root cause"170 command: "kubectl set env deployment/order-service PAYMENT_CIRCUIT_BREAKER=open"171 verify: "5xx rate drops, orders queue for retry"172173 - option: Database failover174 when: "Primary database unresponsive"175 command: "aws rds failover-db-cluster --db-cluster-identifier orders-cluster"176 verify: "Connections re-establish within 60 seconds"177178recovery:179 - "Confirm baseline error rate for 15 minutes"180 - "Check for data inconsistencies from failed transactions"181 - "Update status page to resolved"182 - "Schedule postmortem within 48 hours"183```184185## Escalation Policies186187| Trigger | Action | Timeout |188|---------|--------|---------|189| Alert fires | Page primary on-call | -- |190| No ack in 5 min | Escalate to secondary | 5 min |191| No ack in 15 min | Escalate to engineering manager | 10 min |192| SEV1 declared | Auto-notify VP Eng + CTO | Immediate |193| 30 min without mitigation | IC requests additional responders | IC decision |194| Customer data exposed | Notify Security + Legal | Immediate |195196### Communication Templates197198```markdown199# Status Page -- Investigating200**[Investigating] Elevated error rates on checkout**201We are investigating errors during checkout. Our team is engaged.202Update within 30 minutes.203204# Status Page -- Identified205**[Identified] Checkout errors caused by payment config issue**206Root cause identified. Fix deploying now. Update in 15 minutes.207208# Status Page -- Resolved209**[Resolved] Checkout errors resolved**210Configuration rolled back at 14:47 UTC. All systems normal.211Failed orders auto-retried. Full report within 48 hours.212213# Internal Update (Slack #incidents)214**SEV1 Update -- 14:45 UTC**215Impact: Checkout down since 14:25 | Root cause: bad config deploy216Action: Rollback in progress, ETA 5 min217Revenue impact: ~$12,000 est | IC: @alice218```219220## PagerDuty / OpsGenie Integration221222### PagerDuty Event API v2223224```typescript225async function triggerIncident(params: {226 title: string;227 severity: 'critical' | 'error' | 'warning' | 'info';228 service: string;229 dedupKey: string;230}): Promise<string> {231 const response = await fetch('https://events.pagerduty.com/v2/enqueue', {232 method: 'POST',233 headers: { 'Content-Type': 'application/json' },234 body: JSON.stringify({235 routing_key: process.env.PAGERDUTY_ROUTING_KEY,236 event_action: 'trigger',237 dedup_key: params.dedupKey,238 payload: {239 summary: params.title,240 severity: params.severity,241 source: params.service,242 },243 links: [244 { href: `https://grafana.internal/d/${params.service}`, text: 'Dashboard' },245 { href: `https://wiki.internal/runbooks/${params.service}`, text: 'Runbook' },246 ],247 }),248 });249 return (await response.json()).dedup_key;250}251252async function resolveIncident(dedupKey: string): Promise<void> {253 await fetch('https://events.pagerduty.com/v2/enqueue', {254 method: 'POST',255 headers: { 'Content-Type': 'application/json' },256 body: JSON.stringify({257 routing_key: process.env.PAGERDUTY_ROUTING_KEY,258 event_action: 'resolve',259 dedup_key: dedupKey,260 }),261 });262}263```264265## Postmortem Template266267```markdown268# Postmortem: [Incident Title]269270**Date:** YYYY-MM-DD | **Duration:** HH:MM | **Severity:** SEVN271**IC:** [Name] | **Authors:** [Names] | **Status:** Draft / Complete272273## Summary274One paragraph: what happened, impact, resolution.275276## Impact277| Metric | Value |278|--------|-------|279| User impact duration | X hours Y minutes |280| Users affected | N (Z% of total) |281| Revenue impact | $X,XXX estimated |282| SLA impact | X min against 99.9% target |283284## Timeline (UTC)285| Time | Event |286|------|-------|287| 14:25 | Deploy #4521 pushed (config change) |288| 14:32 | Alert fires: HighErrorBudgetBurnRate_Fast |289| 14:36 | SEV1 declared, war room opened |290| 14:38 | Root cause: payment timeout changed 30s -> 3s |291| 14:42 | Rollback initiated |292| 14:47 | Rollback complete | 14:55 | Fully stable |293294## Root Cause295[What broke and why -- focus on systems, not individuals]296297## Lessons Learned298### What went well299- [Bullet points]300### What went poorly301- [Bullet points]302### Where we got lucky303- [Bullet points]304305## Action Items306| ID | Action | Priority | Owner | Due | Status |307|----|--------|----------|-------|-----|--------|308| 1 | [Action] | P1 | [Name] | [Date] | Open |309```310311## Blameless Postmortem Culture312313| Principle | In Practice |314|-----------|------------|315| People are not the root cause | "The pipeline allowed unsafe config" not "Alice deployed bad config" |316| Focus on systems | Identify process gaps, missing guardrails, tooling deficiencies |317| Assume good intentions | Everyone tried to do the right thing with available info |318| No counterfactuals | Not "if only X..." but "what system change prevents this?" |319| Share widely | Postmortems are learning, not shame |320| Track to completion | Postmortems without follow-through teach nothing |321322### Incident Timeline Reconstruction Sources323324| Source | What It Provides |325|--------|-----------------|326| Alertmanager / PagerDuty | Alert fire/resolve times, ack times, escalations |327| Slack | Human decisions, observations, comms |328| Git / CI | Deploy times, code changes |329| Grafana / Metrics | Anomaly onset, metric correlation |330| Application logs | Error details, trace context |331| Kubernetes events | Pod restarts, OOM kills, scheduling |332| Cloud provider | Infrastructure changes, regional outages |333334## Anti-Patterns335336| Anti-Pattern | Problem | Fix |337|--------------|---------|-----|338| No severity classification | Every incident treated the same | Define and enforce severity matrix |339| Hero culture | One person handles all incidents, burns out | Build rotation with 5+ engineers, document in runbooks |340| Blame-driven postmortems | People hide mistakes, learning stops | Enforce blameless process, focus on systems |341| No runbooks | Responders waste 20+ minutes figuring out what to do | Require runbook link on every alert |342| Postmortems without action items | Same incident recurs | Track items in sprint backlog with owners and deadlines |343| Alert without context | "Check failed" with no links | Include dashboard, runbook, impact in every alert |344| No communication plan | Stakeholders flood war room | Assign comms lead, use status page templates |345| Skipping postmortem for small incidents | Miss patterns that compound | Postmortem SEV1-2, lightweight review SEV3 |346| Testing in prod without rollback plan | "Quick fix" makes things worse | Always have rollback command ready first |347| Ignoring near-misses | Only learning from actual incidents | Track and review near-misses monthly |348| War room without clear roles | Everyone talks, nobody acts | Assign IC, Comms, Ops, Scribe at start |349| Over-classifying severity | Everything is SEV1, diluted response | Calibrate quarterly, push back on inflation |350351## Incident Readiness Checklist352353### Infrastructure354355- [ ] On-call rotation configured with primary and secondary356- [ ] Escalation policy tested end-to-end (alert -> page -> ack -> resolve)357- [ ] PagerDuty / OpsGenie integrated with monitoring stack358- [ ] Incident Slack channel creation automated (bot or `/incident`)359- [ ] Status page configured with component hierarchy360361### Process362363- [ ] Severity classification matrix documented and team-trained364- [ ] War room protocol documented with role descriptions365- [ ] Communication templates ready (status page, internal, customer)366- [ ] Escalation matrix documented (who to call, when)367- [ ] Postmortem template in shared repository368- [ ] Action item tracking integrated with sprint planning369370### Runbooks371372- [ ] Every critical alert has a linked runbook373- [ ] Runbooks include diagnosis steps with actual commands374- [ ] Runbooks include mitigation options with verification375- [ ] Runbooks reviewed and updated quarterly376377### Practice378379- [ ] Game day / incident simulation conducted quarterly380- [ ] New on-call engineers shadow for one rotation first381- [ ] Postmortem review within 48 hours of SEV1/SEV2382- [ ] Monthly review of incident trends (frequency, MTTR, severity)383- [ ] Quarterly review of on-call burden (pages per shift, wake-ups)