War Room Protocol
Phase 1: Activation Criteria
Define when a war room should be activated.
Activate when any of the following are true:
- Customer-facing service is fully down
- Data loss or data integrity issue confirmed
- Security breach in progress
- Revenue impact exceeds $___ / hour
- SLA breach imminent or confirmed
- Multiple teams required for resolution
Activation process:
- On-call engineer escalates to incident commander
- Incident commander declares war room
- War room channel/bridge is created
- Required roles are paged
Phase 2: Role Assignments
| Role | Responsibility | Current Assignee |
|---|---|---|
| Incident Commander (IC) | Coordinates response, makes decisions, manages timeline | |
| Technical Lead | Drives technical investigation and resolution | |
| Communications Lead | Manages stakeholder updates (internal and external) | |
| Scribe | Documents timeline, actions, and decisions | |
| Subject Matter Experts | Provide domain expertise as needed |
Role Rules:
- IC does NOT debug — they coordinate
- One person per role (no shared responsibilities)
- Roles can be handed off with explicit verbal acknowledgment
- IC can request any engineer join the war room
Phase 3: Communication Cadence
| Audience | Channel | Frequency | Owner |
|---|---|---|---|
| War room participants | Voice bridge + chat channel | Continuous | IC |
| Engineering leadership | Status update | Every 30 min | Communications Lead |
| Customer support | Status page + internal brief | Every 30 min | Communications Lead |
| Affected customers | Status page / email | Every 60 min or on status change | Communications Lead |
| Executive team | Summary brief | Every 60 min | IC |
Update Template:
Status: Investigating / Identified / Monitoring / Resolved
Impact: [description of customer impact]
Current action: [what is being done right now]
Next update: [time]
Phase 4: Resolution Workflow
Assess (first 15 minutes)
- Confirm and quantify impact
- Identify affected systems and services
- Review recent changes (deploys, config changes, infra changes)
Stabilize (parallel workstreams)
- Attempt rollback of recent changes if applicable
- Apply mitigation (failover, traffic shift, feature disable)
- Scale resources if capacity-related
Resolve
- Identify root cause
- Apply fix
- Verify fix in production
- Monitor for recurrence (minimum 30 minutes)
Close
- IC declares incident resolved
- Send final communication to all audiences
- Schedule postmortem within 48 hours
- Create follow-up tickets for permanent fixes
Phase 5: War Room Etiquette
- Keep the voice bridge clear — use chat for non-urgent items
- Prefix messages:
[UPDATE],[QUESTION],[ACTION],[FYI] - No blame — focus on resolution
- All decisions and actions are logged by the scribe
- Anyone can call a timeout if the approach is not working
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |
Output Format
Summary
- Incident: ___
- Severity: ___
- Duration: ___
- Impact: ___
- Root cause: ___
- Resolution: ___
Action Items
- Conduct postmortem within 48 hours
- File tickets for all follow-up remediation items
- Update runbooks based on lessons learned
- Review and update war room protocol if gaps were found
- Recognize team members who contributed to resolution