Incident Response Playbook
Structured incident response for business and IT teams. Guides you through detection, triage, containment, resolution, and post-mortem — with auto-generated timelines and action items.
What It Does
When triggered with an incident description, this skill:
- Classifies severity (P1-P4) based on impact and urgency
- Generates a response checklist tailored to incident type (outage, data breach, security event, service degradation, vendor failure)
- Builds a communication plan — who to notify, when, what channels
- Creates a real-time timeline as you log updates
- Produces a post-mortem template with root cause analysis and prevention steps
Usage
Tell your agent about an incident:
"Production API is returning 500 errors for 20% of requests. Started 10 minutes ago."
Or trigger proactively:
"Create an incident response plan for a potential data breach scenario"
Incident Types Covered
- Service outages — full or partial downtime
- Security incidents — breaches, unauthorized access, phishing
- Data incidents — corruption, loss, privacy violations
- Vendor failures — third-party SLA breaches
- Performance degradation — latency spikes, capacity issues
Severity Matrix
| Level |
Impact |
Response Time |
Escalation |
| P1 - Critical |
Business stopped |
Immediate |
Executive + all hands |
| P2 - High |
Major feature down |
< 30 min |
Engineering lead + PM |
| P3 - Medium |
Degraded experience |
< 2 hours |
On-call team |
| P4 - Low |
Minor issue |
Next business day |
Ticket queue |
Response Framework
1. Detection & Triage (First 5 minutes)
- Confirm the incident is real (not a false alarm)
- Classify severity using the matrix above
- Assign incident commander
- Open a dedicated communication channel
2. Containment (First 30 minutes)
- Identify blast radius — what's affected?
- Apply immediate mitigation (rollback, feature flag, scaling)
- Communicate status to stakeholders
3. Resolution
- Root cause investigation
- Implement fix with verification
- Monitor for recurrence
- Update all stakeholders
4. Post-Mortem (Within 48 hours)
- Timeline of events
- Root cause analysis (5 Whys)
- What went well / what didn't
- Action items with owners and deadlines
- Process improvements
Integration
Works with any monitoring stack. Feed alerts from PagerDuty, Datadog, Grafana, or manual reports.
Pro Tip
Pair this with a full AI Operations Context Pack for your industry. Pre-built incident taxonomies, compliance-aware escalation paths, and automated stakeholder templates.
Browse packs: https://afrexai-cto.github.io/context-packs/
Free tools:
1---2name: incident-response-playbook3description: Structured incident response for business and IT teams. Guides you through detection, triage, containment, resolution, and post-mortem — with auto-generated timelines and action items.4---5
6# Incident Response Playbook
7
8Structured incident response for business and IT teams. Guides you through detection, triage, containment, resolution, and post-mortem — with auto-generated timelines and action items.
9
10## What It Does
11
12When triggered with an incident description, this skill:
13
141. **Classifies severity** (P1-P4) based on impact and urgency
152. **Generates a response checklist** tailored to incident type (outage, data breach, security event, service degradation, vendor failure)
163. **Builds a communication plan** — who to notify, when, what channels
174. **Creates a real-time timeline** as you log updates
185. **Produces a post-mortem template** with root cause analysis and prevention steps
19
20## Usage
21
22Tell your agent about an incident:
23
24> "Production API is returning 500 errors for 20% of requests. Started 10 minutes ago."
25
26Or trigger proactively:
27
28> "Create an incident response plan for a potential data breach scenario"
29
30## Incident Types Covered
31
32- **Service outages** — full or partial downtime
33- **Security incidents** — breaches, unauthorized access, phishing
34- **Data incidents** — corruption, loss, privacy violations
35- **Vendor failures** — third-party SLA breaches
36- **Performance degradation** — latency spikes, capacity issues
37
38## Severity Matrix
39
40| Level | Impact | Response Time | Escalation |
41|-------|--------|---------------|------------|
42| P1 - Critical | Business stopped | Immediate | Executive + all hands |
43| P2 - High | Major feature down | < 30 min | Engineering lead + PM |
44| P3 - Medium | Degraded experience | < 2 hours | On-call team |
45| P4 - Low | Minor issue | Next business day | Ticket queue |
46
47## Response Framework
48
49### 1. Detection & Triage (First 5 minutes)
50- Confirm the incident is real (not a false alarm)
51- Classify severity using the matrix above
52- Assign incident commander
53- Open a dedicated communication channel
54
55### 2. Containment (First 30 minutes)
56- Identify blast radius — what's affected?
57- Apply immediate mitigation (rollback, feature flag, scaling)
58- Communicate status to stakeholders
59
60### 3. Resolution
61- Root cause investigation
62- Implement fix with verification
63- Monitor for recurrence
64- Update all stakeholders
65
66### 4. Post-Mortem (Within 48 hours)
67- Timeline of events
68- Root cause analysis (5 Whys)
69- What went well / what didn't
70- Action items with owners and deadlines
71- Process improvements
72
73## Integration
74
75Works with any monitoring stack. Feed alerts from PagerDuty, Datadog, Grafana, or manual reports.
76
77## Pro Tip
78
79Pair this with a full **AI Operations Context Pack** for your industry. Pre-built incident taxonomies, compliance-aware escalation paths, and automated stakeholder templates.
80
81Browse packs: https://afrexai-cto.github.io/context-packs/
82
83Free tools:
84- AI Revenue Calculator: https://afrexai-cto.github.io/ai-revenue-calculator/
85- Agent Setup Wizard: https://afrexai-cto.github.io/agent-setup/