Slo And Alerting

Engineering service-level contracts and actionable alerting: defining user-centered SLIs and SLOs with explicit populations and windows, negotiating error-budget policy, separating request- and time-based semantics, deriving multi-window burn alerts, handling low traffic and missing data, and routing symptoms or predictive hazards by urgency and actionability. Use when an SLO is ambiguous, an SLA lacks operating margin, alert noise is high, resource thresholds page without context, or PromQL burn rules need review. Instrument design belongs to metrics-and-cardinality; percentile semantics to latency-statistics; overload controls to rate-limiting-and-load-shedding.

robsonkades Updated

File contents

robsonkades/agent-skills/tree/main/skills/slo-and-alerting commit 65a49a8b4f

Frequently asked questions

npx skillmds@latest add robsonkades/slo-and-alerting