SRE Dashboards
Build dashboards that help teams detect, triage, and prevent reliability incidents.
When to Use This Skill
Use this skill when:
- Defining service-level dashboards for production systems
- Tracking SLO health and error-budget burn
- Creating incident command-center views
- Standardizing dashboard patterns across teams
Prerequisites
- Metrics pipeline (Prometheus, OpenTelemetry, or vendor equivalent)
- Logs/traces linked to services and environments
- Agreed service taxonomy (team, service, tier, environment)
Dashboard Architecture
Structure dashboards in layers:
- Executive Reliability View: SLO attainment, incident counts, MTTR trends.
- Service Health View: RED/USE metrics, dependency health, release markers.
- Deep-Dive View: Per-endpoint latency, resource saturation, error categories.
Keep each view answer-oriented:
- Are customers impacted?
- What changed?
- Where is the bottleneck?
Core SRE Panels
Golden Signals
- Latency: p50/p95/p99 request duration by endpoint
- Traffic: request throughput and queue depth
- Errors: 5xx rate, failed jobs, timeout ratio
- Saturation: CPU, memory, disk I/O, thread/connection pool exhaustion
SLO Panels
- Current SLI value (rolling windows: 5m, 1h, 24h, 30d)
- Error-budget remaining (%)
- Burn-rate panels (fast and slow windows)
- Multi-window burn alert status
Change Correlation
- Deployment markers and config-change annotations
- Feature flag state overlays
- Upstream/downstream dependency error rates
Example PromQL Snippets
# API error rate (%)
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# p95 latency by route
histogram_quantile(0.95,
sum by (le, route) (rate(http_request_duration_seconds_bucket[5m]))
)
# Fast burn rate (5m / 1h)
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
)
/
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/ sum(rate(http_requests_total[1h]))
)
Operational Guidelines
- Use consistent color semantics (green=healthy, yellow=degrading, red=breach)
- Label units explicitly (ms, req/s, %, cores)
- Default time windows to incident-friendly ranges (15m, 1h, 6h, 24h)
- Minimize panel count per dashboard to reduce cognitive load
- Add runbook links directly in panel descriptions
Troubleshooting
Panel appears flat or empty
- Verify label cardinality and filters (
service, env, region)
- Confirm scrape/ingest latency is within expected range
- Check metric rename regressions after instrumentation updates
High cardinality slows dashboards
- Aggregate by stable dimensions (
service, route_group) instead of raw IDs
- Use recording rules for expensive percentile and ratio queries
- Split deep-dive dashboards from NOC summary dashboards
Related Skills
1---2name: sre-dashboards3description: Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. Use when building observability views for SLOs, incident response, and executive reliability reporting.4license: MIT5---6
7# SRE Dashboards
8
9Build dashboards that help teams detect, triage, and prevent reliability incidents.
10
11## When to Use This Skill
12
13Use this skill when:
14- Defining service-level dashboards for production systems
15- Tracking SLO health and error-budget burn
16- Creating incident command-center views
17- Standardizing dashboard patterns across teams
18
19## Prerequisites
20
21- Metrics pipeline (Prometheus, OpenTelemetry, or vendor equivalent)
22- Logs/traces linked to services and environments
23- Agreed service taxonomy (team, service, tier, environment)
24
25## Dashboard Architecture
26
27Structure dashboards in layers:
28
291. **Executive Reliability View**: SLO attainment, incident counts, MTTR trends.
302. **Service Health View**: RED/USE metrics, dependency health, release markers.
313. **Deep-Dive View**: Per-endpoint latency, resource saturation, error categories.
32
33Keep each view answer-oriented:
34- *Are customers impacted?*
35- *What changed?*
36- *Where is the bottleneck?*
37
38## Core SRE Panels
39
40### Golden Signals
41
42- **Latency**: p50/p95/p99 request duration by endpoint
43- **Traffic**: request throughput and queue depth
44- **Errors**: 5xx rate, failed jobs, timeout ratio
45- **Saturation**: CPU, memory, disk I/O, thread/connection pool exhaustion
46
47### SLO Panels
48
49- Current SLI value (rolling windows: 5m, 1h, 24h, 30d)
50- Error-budget remaining (%)
51- Burn-rate panels (fast and slow windows)
52- Multi-window burn alert status
53
54### Change Correlation
55
56- Deployment markers and config-change annotations
57- Feature flag state overlays
58- Upstream/downstream dependency error rates
59
60## Example PromQL Snippets
61
62```promql
63# API error rate (%)
64100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
65 / sum(rate(http_requests_total[5m]))
66```
67
68```promql
69# p95 latency by route
70histogram_quantile(0.95,
71 sum by (le, route) (rate(http_request_duration_seconds_bucket[5m]))
72)
73```
74
75```promql
76# Fast burn rate (5m / 1h)
77(
78 sum(rate(http_requests_total{status=~"5.."}[5m]))
79 / sum(rate(http_requests_total[5m]))
80)
81/
82(
83 sum(rate(http_requests_total{status=~"5.."}[1h]))
84 / sum(rate(http_requests_total[1h]))
85)
86```
87
88## Operational Guidelines
89
90- Use consistent color semantics (green=healthy, yellow=degrading, red=breach)
91- Label units explicitly (ms, req/s, %, cores)
92- Default time windows to incident-friendly ranges (15m, 1h, 6h, 24h)
93- Minimize panel count per dashboard to reduce cognitive load
94- Add runbook links directly in panel descriptions
95
96## Troubleshooting
97
98### Panel appears flat or empty
99
100- Verify label cardinality and filters (`service`, `env`, `region`)
101- Confirm scrape/ingest latency is within expected range
102- Check metric rename regressions after instrumentation updates
103
104### High cardinality slows dashboards
105
106- Aggregate by stable dimensions (`service`, `route_group`) instead of raw IDs
107- Use recording rules for expensive percentile and ratio queries
108- Split deep-dive dashboards from NOC summary dashboards
109
110## Related Skills
111
112- [prometheus-grafana](../prometheus-grafana/) - Dashboard implementation and PromQL
113- [opentelemetry](../opentelemetry/) - Standardized telemetry instrumentation
114- [alerting-oncall](../alerting-oncall/) - Reliability alert routing and escalation
115- [agent-observability](../../ai/agent-observability/) - AI workload reliability telemetry