Incident Metrics Review
Team: {{ team_name }} | Period: {{ review_period }} Data Source: {{ incident_source }}
Core Metrics Definitions
MTTD — Mean Time to Detect
Time from when an incident begins to when it is detected by monitoring or reported.
Formula: MTTD = Average(detection_time - incident_start_time)
Why it matters: Long MTTD means incidents are silently impacting users before anyone knows.
MTTA — Mean Time to Acknowledge
Time from when an alert fires to when a responder acknowledges it.
Formula: MTTA = Average(acknowledge_time - alert_fire_time)
Why it matters: High MTTA indicates on-call responsiveness issues or alert fatigue.
MTTR — Mean Time to Resolve
Time from incident detection to full resolution.
Formula: MTTR = Average(resolution_time - detection_time)
Why it matters: Primary measure of incident response effectiveness.
MTBF — Mean Time Between Failures
Time between the resolution of one incident and the start of the next.
Formula: MTBF = Average(next_incident_start - previous_incident_resolution)
Why it matters: Low MTBF indicates systemic reliability issues.
Data Collection Template
Incident Log for {{ review_period }}
| # | Incident | Severity | Start Time | Detected | Acknowledged | Resolved | MTTD | MTTA | MTTR |
|---|---|---|---|---|---|---|---|---|---|
| 1 | title | SEV | time | time | time | time | min | min | min |
| 2 | title | SEV | time | time | time | time | min | min | min |
Summary Statistics
| Metric | SEV1 | SEV2 | SEV3 | All |
|---|---|---|---|---|
| Count | n | n | n | n |
| MTTD (avg) | min | min | min | min |
| MTTA (avg) | min | min | min | min |
| MTTR (avg) | min | min | min | min |
| MTBF (avg) | days | days | days | days |
Benchmarks
Industry benchmarks for reference (adjust based on your context):
| Metric | Excellent | Good | Needs Improvement |
|---|---|---|---|
| MTTD | < 5 min | 5-15 min | > 15 min |
| MTTA | < 5 min | 5-15 min | > 15 min |
| MTTR (SEV1) | < 30 min | 30-60 min | > 60 min |
| MTTR (SEV2) | < 2 hours | 2-4 hours | > 4 hours |
| MTBF | > 30 days | 14-30 days | < 14 days |
Trend Analysis
Month-over-Month Comparison
| Metric | Previous Period | Current Period | Change | Trend |
|---|---|---|---|---|
| Incident Count | n | n | +/- | improving/stable/declining |
| MTTD | min | min | +/- | improving/stable/declining |
| MTTA | min | min | +/- | improving/stable/declining |
| MTTR | min | min | +/- | improving/stable/declining |
| MTBF | days | days | +/- | improving/stable/declining |
Distribution Analysis
- Incidents by severity: % SEV1, % SEV2, % SEV3, % SEV4
- Incidents by time of day: % business hours, % after-hours, % weekends
- Incidents by service: top 3 services by incident count
- Incidents by root cause category: deployment, infrastructure, dependency, etc.
Improvement Recommendations
If MTTD is High
- Implement synthetic monitoring for critical user journeys
- Add anomaly detection to key business metrics
- Review alert coverage for gaps in observability
- Consider real-user monitoring (RUM) for client-side detection
If MTTA is High
- Review on-call notification channels (push vs. SMS vs. phone)
- Audit alert routing rules for correctness
- Address alert fatigue by reducing false positives
- Review on-call engineer workload and burnout indicators
If MTTR is High
- Invest in runbook automation for common incident types
- Improve diagnostic tooling and dashboards
- Conduct game days to practice incident response
- Review escalation policies for faster expert engagement
- Pre-build rollback procedures for every deployment
If MTBF is Low
- Focus on systemic reliability improvements (redundancy, resilience)
- Review and prioritize postmortem action items
- Invest in chaos engineering to proactively find weaknesses
- Increase test coverage for failure scenarios
Action Items
| Action | Impact on Metric | Effort | Owner | Target Date |
|---|---|---|---|---|
| action | MTTD/MTTA/MTTR/MTBF | low/med/high | name | date |
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |