Runbook

Keep a service so boring nobody notices it's there — author and maintain operational runbooks, tune alerts so they fire on real user pain and nothing else, and harden uptime *before* the 3am page instead of after. Use this when writing or reviewing a runbook for a service or a recurring failure, designing or pruning alerts (killing noisy/flappy alerts, fixing alert fatigue, mapping every alert to an action), setting SLOs and error budgets, deciding what's worth paging a human for, prepping or cleaning up an on-call rotation, hardening a deploy or a fragile dependency, or answering "how do we stop getting paged for this." This is the proactive, prevention side of reliability. For commanding a live incident that is already on fire, use `incident-response` instead.

5dive-ai a4e6c5b 2 files · 7.6 KB Updated

File contents

5dive-ai/character-packs/tree/main/packs/anton/skills/runbook commit a4e6c5bbc7

Frequently asked questions

npx skillmds@latest add 5dive-ai/runbook