Incident Response
- Establish impact, start time, affected components, and current state.
- Preserve logs and state before restarting or mutating systems.
- Apply the smallest reversible mitigation that restores service.
- Trace the causal chain with timestamps and correlation identifiers.
- Verify recovery through user-visible health and backlog drain, not process existence alone.
- Document root cause, contributing conditions, detection gaps, and concrete prevention items.