1---2name: observability-maturity-check3description: Use when performing observability maturity check — evaluates an organization's observability maturity across the three pillars (metrics, logs, traces) plus alerting, dashboards, and AIOps capabilities. Identifies gaps in visibility, assesses signal quality, and produces a roadmap for achieving full-stack observability.4---56# Observability Maturity Check78## Phase 1: Metrics Assessment91. Evaluate metrics collection and usage10 - [ ] Infrastructure metrics (CPU, memory, disk, network)11 - [ ] Application metrics (request rate, error rate, latency - RED)12 - [ ] Business metrics (transactions, revenue, user activity)13 - [ ] Custom metrics for domain-specific KPIs14 - [ ] Metrics retention and resolution appropriate15 - [ ] Metrics naming conventions standardized16 - [ ] Service-level indicators (SLIs) derived from metrics172. Score: 1 (Basic infra only) to 5 (Full-stack with business metrics)1819### Metrics Coverage2021| Layer | Coverage | Gaps | Quality |22|-------|----------|------|---------|23| Infrastructure | % of hosts/containers | | High/Med/Low |24| Application (RED) | % of services | | |25| Database | % of instances | | |26| External dependencies | % of integrations | | |27| Business KPIs | % defined | | |2829## Phase 2: Logging Assessment301. Evaluate logging practices31 - [ ] Centralized log aggregation32 - [ ] Structured logging (JSON) across all services33 - [ ] Consistent log levels (DEBUG, INFO, WARN, ERROR)34 - [ ] Request/correlation IDs in all log entries35 - [ ] PII/sensitive data redaction in logs36 - [ ] Log retention policies defined37 - [ ] Log search performance adequate38 - [ ] Log-based alerting configured392. Score: 1 (Scattered file logs) to 5 (Centralized, structured, searchable)4041## Phase 3: Distributed Tracing Assessment421. Evaluate tracing implementation43 - [ ] Tracing instrumented across services44 - [ ] Trace propagation across service boundaries45 - [ ] Span attributes include relevant context46 - [ ] Sampling strategy defined (head, tail, adaptive)47 - [ ] Trace-to-log and trace-to-metric correlation48 - [ ] Service dependency map generated from traces49 - [ ] Trace data used in incident investigation502. Score: 1 (No tracing) to 5 (Full distributed tracing with correlation)5152### Three Pillars Coverage5354| Pillar | Coverage (%) | Quality (1-5) | Tool | Key Gap |55|--------|-------------|--------------|------|---------|56| Metrics | % | | | |57| Logs | % | | | |58| Traces | % | | | |59| **Combined** | **%** | **/5** | | |6061## Phase 4: Alerting Quality Assessment621. Evaluate alerting effectiveness63 - [ ] Alerts are actionable (every alert requires human action)64 - [ ] Alert severity levels match impact65 - [ ] On-call rotation and escalation configured66 - [ ] Alert noise ratio acceptable (signal vs. noise)67 - [ ] Symptom-based alerts (not cause-based)68 - [ ] Runbooks linked to alerts69 - [ ] Alert fatigue measured and managed70 - [ ] SLO-based alerting (burn rate alerts)712. Measure alert quality metrics7273### Alert Quality Metrics7475| Metric | Current | Target | Status |76|--------|---------|--------|--------|77| Alerts/week | | < | |78| Actionable alerts % | % | > 80% | |79| Alerts with runbooks % | % | 100% | |80| MTTA (time to acknowledge) | min | < 5 min | |81| False positive rate | % | < 10% | |82| Duplicate/correlated alerts | % | < 5% | |8384## Phase 5: Dashboards & Visualization851. Evaluate dashboard practices86 - [ ] Service-level dashboards for each team87 - [ ] On-call dashboard (single pane of glass)88 - [ ] SLO status dashboards89 - [ ] Infrastructure overview dashboards90 - [ ] Business metrics dashboards91 - [ ] Dashboard naming and organization standards92 - [ ] Dashboard-as-code (version controlled)932. Assess dashboard discoverability and usefulness9495## Phase 6: Advanced Capabilities961. Evaluate advanced observability features97 - [ ] Anomaly detection (automated baseline comparison)98 - [ ] Root cause analysis automation99 - [ ] Service dependency visualization100 - [ ] Change correlation (deploy/config changes vs. incidents)101 - [ ] Continuous profiling (CPU, memory, lock)102 - [ ] Real User Monitoring (RUM) / frontend observability103 - [ ] Synthetic monitoring for critical paths104 - [ ] OpenTelemetry adoption for vendor neutrality105106### Maturity Scorecard107108| Dimension | Score (1-5) | Current State | Target State |109|-----------|-----------|---------------|-------------|110| Metrics | | | |111| Logging | | | |112| Tracing | | | |113| Alerting | | | |114| Dashboards | | | |115| Advanced (AI/ML) | | | |116| **Overall** | **/5** | | |117118## Counter-Rationalizations119120| Shortcut | Counter | Why |121|----------|---------|-----|122| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |123| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |124| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |125| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |126| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |127128## Output Format129- **Pillar Assessment**: Metrics, logs, traces coverage and quality130- **Alert Quality Report**: Noise analysis and improvement recommendations131- **Coverage Gaps**: Services and layers lacking observability132- **Tool Assessment**: Current stack evaluation and recommendations133- **Improvement Roadmap**: Phased plan to advance observability maturity134135## Action Items136- [ ] Assess coverage across all three pillars137- [ ] Audit alert quality and reduce noise138- [ ] Implement structured logging across all services139- [ ] Roll out distributed tracing to uninstrumented services140- [ ] Standardize dashboards and make discoverable141- [ ] Evaluate and pilot advanced capabilities142- [ ] Schedule quarterly observability review