Incident Response

Stabilize a system that is failing right now — name the signal that flagged it, size the blast radius in numbers, keep a timestamped log written as you go, find the last known-good state, then propose the smallest reversible mitigation and confirm recovery against that same signal. Use when production is degraded or down and time to mitigation matters more than a complete explanation. Not for a defect that is not currently failing (systematic-debugging), not for the root-cause investigation or the postmortem that follows, and never authorization to touch production without asking first.

nahid-sparktales Updated

File contents

nahid-sparktales/agent-dispatcher/tree/main/skills/devops/incident-response commit f1dd06e735

Frequently asked questions

npx skillmds@latest add nahid-sparktales/incident-response