Incident Timeline Reconstruction
Incident: {{ incident_title }}
Window: {{ incident_start }} to {{ incident_end }}
Purpose
A precise, evidence-backed timeline is the foundation of every good postmortem. This skill guides you through gathering data from multiple sources, correlating events, and producing a single authoritative timeline.
Data Sources to Collect
1. Monitoring and Alerting
2. Deployment and Change Records
3. Communication Records
4. System Logs
5. Customer Reports
Timeline Template
Record every event with its exact timestamp, source, and type:
| Time (UTC) |
Source |
Type |
Event Description |
| HH:MM:SS |
monitoring/deploy/human/log |
trigger/detection/action/decision/resolution |
what happened |
Event Types
- Trigger — the root cause event (deployment, config change, external failure)
- Impact Start — when users began experiencing the issue
- Detection — when the team became aware (alert, customer report)
- Declaration — when the incident was formally declared
- Investigation — diagnostic actions taken
- Decision — key decisions made by IC
- Action — mitigation or remediation actions executed
- Mitigation — when user impact was reduced or eliminated
- Resolution — when the incident was fully resolved
Correlation Techniques
Clock Synchronization
- Ensure all timestamps are in the same timezone (UTC preferred)
- Account for clock skew between systems (check NTP status)
- Note: Slack timestamps use the sender's timezone by default
Gap Analysis
After building the initial timeline, look for:
- Gaps longer than 5 minutes during active investigation
- Events without a clear causal connection to adjacent events
- Missing human actions (who decided what and when?)
- Discrepancies between verbal accounts and logged evidence
Causal Chain Mapping
For each event, ask:
- What caused this event?
- What did this event cause?
- Was this event preventable?
- Was this event detectable earlier?
Timeline Review Process
- Individual accounts — each responder writes their own timeline from memory
- Evidence merge — combine individual accounts with system evidence
- Group review — walk through the merged timeline as a group
- Gap filling — investigate and resolve discrepancies
- Final timeline — produce the authoritative version for the postmortem
Counter-Rationalizations
| Shortcut |
Counter |
Why |
| "We can skip some steps for this case" |
Adapt the workflow steps, don't skip them |
Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" |
Complete all workflow phases with the user |
The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" |
Scale the process down, don't turn it off |
Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" |
Complete each section before moving on |
Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" |
Always produce the structured output format |
Structured output enables comparison, audit trails, and handoff to other teams |
Output Format
The final timeline should include:
- Every event with UTC timestamp and source
- Clear marking of trigger, detection, and resolution points
- Duration between key phases (trigger→detection, detection→mitigation, mitigation→resolution)
- Annotations for decisions and their rationale
- Gaps explicitly noted as "no data available" rather than omitted
1---2name: incident-timeline-reconstruction3description: Use when performing incident timeline reconstruction — post-incident timeline building framework for reconstructing the sequence of events from logs, alerts, chat messages, deployment records, and monitoring data. Provides structured approaches to gathering evidence, correlating timestamps, identifying gaps, and producing an authoritative incident timeline for postmortem analysis.4---56# Incident Timeline Reconstruction78Incident: **{{ incident_title }}**9Window: **{{ incident_start }}** to **{{ incident_end }}**1011## Purpose1213A precise, evidence-backed timeline is the foundation of every good postmortem. This skill guides you through gathering data from multiple sources, correlating events, and producing a single authoritative timeline.1415## Data Sources to Collect1617### 1. Monitoring and Alerting18- [ ] Alert firing times from PagerDuty/OpsGenie/monitoring tool19- [ ] Dashboard screenshots at key moments20- [ ] Metric anomalies (latency spikes, error rate jumps, traffic drops)21- [ ] Health check failures and recovery times22- [ ] SLO/SLI breach timestamps2324### 2. Deployment and Change Records25- [ ] CI/CD pipeline executions around the incident window26- [ ] Git commits and merges in the 24 hours before the incident27- [ ] Infrastructure changes (Terraform, CloudFormation, Kubernetes applies)28- [ ] Feature flag changes29- [ ] Database migrations30- [ ] Config changes (environment variables, secrets rotation)3132### 3. Communication Records33- [ ] Incident Slack channel messages (with timestamps)34- [ ] Bridge call notes and recordings35- [ ] Email threads related to the incident36- [ ] Status page updates and times3738### 4. System Logs39- [ ] Application logs around the incident window40- [ ] Infrastructure logs (load balancer, container orchestrator)41- [ ] Database slow query logs and error logs42- [ ] Network/firewall logs if relevant43- [ ] Cloud provider event logs (CloudTrail, Activity Log)4445### 5. Customer Reports46- [ ] Support ticket timestamps47- [ ] Social media reports48- [ ] Customer-reported symptoms and timing4950## Timeline Template5152Record every event with its exact timestamp, source, and type:5354| Time (UTC) | Source | Type | Event Description |55|------------|--------|------|-------------------|56| _HH:MM:SS_ | _monitoring/deploy/human/log_ | _trigger/detection/action/decision/resolution_ | _what happened_ |5758### Event Types5960- **Trigger** — the root cause event (deployment, config change, external failure)61- **Impact Start** — when users began experiencing the issue62- **Detection** — when the team became aware (alert, customer report)63- **Declaration** — when the incident was formally declared64- **Investigation** — diagnostic actions taken65- **Decision** — key decisions made by IC66- **Action** — mitigation or remediation actions executed67- **Mitigation** — when user impact was reduced or eliminated68- **Resolution** — when the incident was fully resolved6970## Correlation Techniques7172### Clock Synchronization73- Ensure all timestamps are in the same timezone (UTC preferred)74- Account for clock skew between systems (check NTP status)75- Note: Slack timestamps use the sender's timezone by default7677### Gap Analysis78After building the initial timeline, look for:79- Gaps longer than 5 minutes during active investigation80- Events without a clear causal connection to adjacent events81- Missing human actions (who decided what and when?)82- Discrepancies between verbal accounts and logged evidence8384### Causal Chain Mapping85For each event, ask:86- What caused this event?87- What did this event cause?88- Was this event preventable?89- Was this event detectable earlier?9091## Timeline Review Process92931. **Individual accounts** — each responder writes their own timeline from memory942. **Evidence merge** — combine individual accounts with system evidence953. **Group review** — walk through the merged timeline as a group964. **Gap filling** — investigate and resolve discrepancies975. **Final timeline** — produce the authoritative version for the postmortem9899## Counter-Rationalizations100101| Shortcut | Counter | Why |102|----------|---------|-----|103| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |104| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |105| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |106| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |107| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |108109## Output Format110111The final timeline should include:112- Every event with UTC timestamp and source113- Clear marking of trigger, detection, and resolution points114- Duration between key phases (trigger→detection, detection→mitigation, mitigation→resolution)115- Annotations for decisions and their rationale116- Gaps explicitly noted as "no data available" rather than omitted