Observability Audit

Runs a structured, evidence-based audit of a system's observability — logging, metrics, tracing, alerting, dashboards, SLOs/error budgets, runbooks/on-call readiness, health checks, and log/metric retention & cost — then reports the results as one table (check, area, status, evidence, recommendation). Covers structured logging and log-level discipline, correlation/request/trace IDs threaded through a request's lifecycle, golden-signal (RED/USE) metrics coverage for request-driven paths and background/async jobs, distributed trace-context propagation across service and queue boundaries plus sampling strategy, dashboard existence and content (golden signals vs. raw infra graphs, a clear "is the system healthy" entry point), alert quality (symptom-based vs. cause-based, alert fatigue, ownership/runbook links), SLO/error-budget definition and whether it's measured against real production data, on-call runbook coverage and escalation-path documentation, liveness/readiness health-check depth and whether they're act

finnley07 e41a39a 20.5 KB Updated

File contents

finnley07/AI-SKILLHUB/tree/main/observability-audit commit e41a39ac02

Frequently asked questions

npx skillmds@latest add finnley07/observability-audit