Chaos Engineering Incident Drill
Drill: {{ drill_name }} Target: {{ target_service }} | Environment: {{ environment }} Scheduled: {{ drill_date }}
Pre-Drill Planning
Hypothesis
Define what you expect to happen:
Hypothesis: When [failure injection], we expect [expected behavior]. The system should [recover/failover/degrade gracefully] within [time threshold] and monitoring should detect the issue within [detection threshold].
Blast Radius Controls
- Drill is scoped to a single service / availability zone / percentage of traffic
- Kill switch is prepared and tested (how to stop the experiment instantly)
- Affected downstream services identified and owners notified
- Customer impact assessment completed (expected: none / minimal / controlled)
- Rollback procedure documented and ready
Prerequisites Checklist
- Stakeholders notified (engineering, SRE, support, management)
- On-call team aware and standing by
- Monitoring dashboards open and shared
- Baseline metrics captured (latency, error rate, throughput)
- Experiment tooling tested (Chaos Monkey, Litmus, Gremlin, etc.)
- Incident channel created for drill coordination
- No conflicting deployments or maintenance windows
- Runbooks for the target service reviewed
Abort Criteria
Stop the experiment immediately if:
- Customer impact exceeds expected threshold
- Error rates exceed ___% for more than ___ minutes
- Cascading failures detected beyond target blast radius
- Recovery does not begin within ___ minutes
- A real incident is declared during the drill
Experiment Design
Failure Injection Scenarios
| Scenario | Injection Method | Expected Behavior | Duration |
|---|---|---|---|
| e.g., Kill primary DB | terminate instance | failover to replica | 5 min |
| e.g., Network partition | block traffic on port | circuit breaker activates | 3 min |
| e.g., CPU saturation | stress-ng 100% CPU | autoscaling triggers | 10 min |
Observation Points
| What to Observe | Dashboard/Tool | Baseline Value | Threshold |
|---|---|---|---|
| Error rate | Grafana/Datadog | < 0.1% | > 1% |
| Latency p99 | APM tool | 200ms | > 2000ms |
| Alert firing | PagerDuty | none | within 5 min |
| Auto-recovery | K8s/ASG | N/A | within 10 min |
Drill Execution
Phase 1: Baseline (T-10 min)
- Capture baseline metrics screenshot
- Confirm all monitoring is green
- Announce drill start in incident channel
- Confirm kill switch operator is ready
Phase 2: Injection (T=0)
- Execute failure injection
- Start timer
- Begin observation log
Phase 3: Observation (T+0 to T+N)
Observation Log:
| Time | Observation | Expected? | Notes |
|---|---|---|---|
| T+0 | injection executed | yes | — |
| T+1m | what happened | yes/no | — |
| T+5m | what happened | yes/no | — |
Phase 4: Recovery (after injection ends or abort)
- Stop failure injection / activate kill switch
- Monitor recovery metrics
- Record time to recovery
- Verify service returns to baseline
Phase 5: Wrap-Up
- Announce drill complete in incident channel
- Capture post-drill metrics screenshot
Post-Drill Analysis
Results Summary
| Aspect | Expected | Actual | Pass/Fail |
|---|---|---|---|
| Detection time | X min | Y min | pass/fail |
| Alert fired correctly | yes | yes/no | pass/fail |
| Auto-recovery worked | yes | yes/no | pass/fail |
| Recovery time | X min | Y min | pass/fail |
| Customer impact | none | none/minimal | pass/fail |
| Blast radius contained | yes | yes/no | pass/fail |
Hypothesis Validation
- Confirmed / Partially Confirmed / Refuted
- Key findings: ___
Findings and Action Items
| Finding | Severity | Action Item | Owner | Ticket |
|---|---|---|---|---|
| finding | high/med/low | action | name | link |
Recommendations for Next Drill
- Suggested next experiment: ___
- Environment upgrade needed: ___
- Tooling improvements: ___
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |