Monitoring Stack Deployer
Expert in deploying and configuring production monitoring with Prometheus, Grafana, and SLO-driven alerting.
Activation Triggers
Activate on: "monitoring setup", "Prometheus config", "Grafana dashboard", "alerting rules", "SLO dashboard", "metrics pipeline", "observability stack", "kube-prometheus-stack", "ServiceMonitor"
NOT for: Application logging → log-aggregation-architect | Distributed tracing → logging-observability | Incident response → site-reliability-engineer
Quick Start
- Deploy kube-prometheus-stack — Prometheus, Grafana, Alertmanager, node-exporter in one Helm chart
- Define SLOs — availability and latency targets per service
- Create ServiceMonitors — auto-discover application metrics endpoints
- Build dashboards — USE method (utilization, saturation, errors) for infrastructure; RED method (rate, errors, duration) for services
- Configure alerting — SLO burn-rate alerts, not threshold alerts
Core Capabilities
| Domain |
Technologies |
| Metrics |
Prometheus 3.x, Mimir, Thanos, VictoriaMetrics |
| Visualization |
Grafana 11, Perses (open-source Grafana alternative) |
| Alerting |
Alertmanager, PagerDuty, OpsGenie, Slack integration |
| SLOs |
Sloth, Pyrra, Google SRE workbook burn-rate model |
| K8s Native |
kube-prometheus-stack, ServiceMonitor, PodMonitor, PrometheusRule |
Architecture Patterns
SLO-Based Burn-Rate Alerting
Traditional (BAD): "Alert if error rate > 1% for 5 minutes"
Problem: Too many false positives, alert fatigue
SLO-Based (GOOD): "Alert if burning SLO budget too fast"
SLO: 99.9% availability over 30 days → 43.2 min error budget
Multi-window burn rate:
┌─────────────────────────────────────────────┐
│ Severity │ Burn Rate │ Long Window │ Short │
│ Critical │ 14.4x │ 1 hour │ 5 min │
│ Warning │ 6x │ 6 hours │ 30 min │
│ Ticket │ 1x │ 3 days │ 6 hrs │
└─────────────────────────────────────────────┘
Prometheus Recording Rules for SLOs
# PrometheusRule for SLO burn rate
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: api-slo-rules
spec:
groups:
- name: api-slo-burn-rate
rules:
- record: slo:api_availability:burn_rate_1h
expr: |
1 - (
sum(rate(http_requests_total{code!~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
)
/ (1 - 0.999)
- alert: APIAvailabilityBurnRateCritical
expr: slo:api_availability:burn_rate_1h > 14.4
and slo:api_availability:burn_rate_5m > 14.4
for: 2m
labels:
severity: critical
annotations:
summary: "API burning error budget 14.4x faster than allowed"
RED Method Dashboard Layout
┌─────────────────────────────────────────────────────────┐
│ Service: api-gateway SLO: 99.9% │
├──────────────┬──────────────┬───────────────────────────┤
│ RATE │ ERRORS │ DURATION │
│ req/sec │ error % │ p50 / p95 / p99 │
│ ▁▂▃▅▇█▇▅▃ │ ▁▁▁▂▁▁▁▁▁ │ p50: 12ms │
│ peak: 1.2k │ curr: 0.02% │ p95: 89ms p99: 240ms │
├──────────────┴──────────────┴───────────────────────────┤
│ Error Budget: 38.2 min remaining (88% of 43.2 min) │
│ ████████████████████████████████░░░░ │
└─────────────────────────────────────────────────────────┘
Anti-Patterns
- Threshold-based alerting — static thresholds like "alert if CPU > 80%" cause alert fatigue. Use SLO burn rates that correlate with user impact.
- No recording rules — computing complex queries at alert evaluation time is slow and expensive. Pre-compute with recording rules.
- Dashboard sprawl — hundreds of dashboards nobody checks. Build one service dashboard template, parameterize with variables.
- Missing service discovery — manually listing scrape targets. Use ServiceMonitor/PodMonitor to auto-discover Kubernetes workloads.
- Alerting without runbooks — alerts fire but responders do not know what to do. Every alert must link to a runbook with diagnosis steps.
Quality Checklist
[ ] kube-prometheus-stack or equivalent deployed and healthy
[ ] ServiceMonitors auto-discover all application metrics endpoints
[ ] SLOs defined for every user-facing service
[ ] Burn-rate alerts configured (critical, warning, ticket)
[ ] Recording rules pre-compute expensive queries
[ ] Grafana dashboards use RED method for services, USE for infrastructure
[ ] Alertmanager routes to correct channels (PagerDuty/Slack/OpsGenie)
[ ] Alert grouping and inhibition rules prevent notification storms
[ ] Every alert has a linked runbook
[ ] Metrics retention configured (15d local, long-term in Mimir/Thanos)
[ ] Dashboard provisioned as code (JSON/YAML in Git)
[ ] Error budget dashboard visible to engineering and product
1---2name: monitoring-stack-deployer3description: Production monitoring stack deployer with Prometheus, Grafana, and SLO-based alerting. Activate on: monitoring setup, Prometheus configuration, Grafana dashboards, alerting rules, SLO definition, metrics pipeline, observability stack. NOT for: application logging (use log-aggregation-architect), distributed tracing (use logging-observability), incident response (use site-reliability-engineer).4license: Apache-2.05---6
7# Monitoring Stack Deployer
8
9Expert in deploying and configuring production monitoring with Prometheus, Grafana, and SLO-driven alerting.
10
11## Activation Triggers
12
13**Activate on:** "monitoring setup", "Prometheus config", "Grafana dashboard", "alerting rules", "SLO dashboard", "metrics pipeline", "observability stack", "kube-prometheus-stack", "ServiceMonitor"
14
15**NOT for:** Application logging → `log-aggregation-architect` | Distributed tracing → `logging-observability` | Incident response → `site-reliability-engineer`
16
17## Quick Start
18
191. **Deploy kube-prometheus-stack** — Prometheus, Grafana, Alertmanager, node-exporter in one Helm chart
202. **Define SLOs** — availability and latency targets per service
213. **Create ServiceMonitors** — auto-discover application metrics endpoints
224. **Build dashboards** — USE method (utilization, saturation, errors) for infrastructure; RED method (rate, errors, duration) for services
235. **Configure alerting** — SLO burn-rate alerts, not threshold alerts
24
25## Core Capabilities
26
27| Domain | Technologies |
28|--------|-------------|
29| **Metrics** | Prometheus 3.x, Mimir, Thanos, VictoriaMetrics |
30| **Visualization** | Grafana 11, Perses (open-source Grafana alternative) |
31| **Alerting** | Alertmanager, PagerDuty, OpsGenie, Slack integration |
32| **SLOs** | Sloth, Pyrra, Google SRE workbook burn-rate model |
33| **K8s Native** | kube-prometheus-stack, ServiceMonitor, PodMonitor, PrometheusRule |
34
35## Architecture Patterns
36
37### SLO-Based Burn-Rate Alerting
38
39```
40Traditional (BAD): "Alert if error rate > 1% for 5 minutes"
41 Problem: Too many false positives, alert fatigue
42
43SLO-Based (GOOD): "Alert if burning SLO budget too fast"
44 SLO: 99.9% availability over 30 days → 43.2 min error budget
45
46 Multi-window burn rate:
47 ┌─────────────────────────────────────────────┐
48 │ Severity │ Burn Rate │ Long Window │ Short │
49 │ Critical │ 14.4x │ 1 hour │ 5 min │
50 │ Warning │ 6x │ 6 hours │ 30 min │
51 │ Ticket │ 1x │ 3 days │ 6 hrs │
52 └─────────────────────────────────────────────┘
53```
54
55### Prometheus Recording Rules for SLOs
56
57```yaml
58# PrometheusRule for SLO burn rate
59apiVersion: monitoring.coreos.com/v1
60kind: PrometheusRule
61metadata:
62 name: api-slo-rules
63spec:
64 groups:
65 - name: api-slo-burn-rate
66 rules:
67 - record: slo:api_availability:burn_rate_1h
68 expr: |
69 1 - (
70 sum(rate(http_requests_total{code!~"5.."}[1h]))
71 /
72 sum(rate(http_requests_total[1h]))
73 )
74 / (1 - 0.999)
75 - alert: APIAvailabilityBurnRateCritical
76 expr: slo:api_availability:burn_rate_1h > 14.4
77 and slo:api_availability:burn_rate_5m > 14.4
78 for: 2m
79 labels:
80 severity: critical
81 annotations:
82 summary: "API burning error budget 14.4x faster than allowed"
83```
84
85### RED Method Dashboard Layout
86
87```
88┌─────────────────────────────────────────────────────────┐
89│ Service: api-gateway SLO: 99.9% │
90├──────────────┬──────────────┬───────────────────────────┤
91│ RATE │ ERRORS │ DURATION │
92│ req/sec │ error % │ p50 / p95 / p99 │
93│ ▁▂▃▅▇█▇▅▃ │ ▁▁▁▂▁▁▁▁▁ │ p50: 12ms │
94│ peak: 1.2k │ curr: 0.02% │ p95: 89ms p99: 240ms │
95├──────────────┴──────────────┴───────────────────────────┤
96│ Error Budget: 38.2 min remaining (88% of 43.2 min) │
97│ ████████████████████████████████░░░░ │
98└─────────────────────────────────────────────────────────┘
99```
100
101## Anti-Patterns
102
1031. **Threshold-based alerting** — static thresholds like "alert if CPU > 80%" cause alert fatigue. Use SLO burn rates that correlate with user impact.
1042. **No recording rules** — computing complex queries at alert evaluation time is slow and expensive. Pre-compute with recording rules.
1053. **Dashboard sprawl** — hundreds of dashboards nobody checks. Build one service dashboard template, parameterize with variables.
1064. **Missing service discovery** — manually listing scrape targets. Use ServiceMonitor/PodMonitor to auto-discover Kubernetes workloads.
1075. **Alerting without runbooks** — alerts fire but responders do not know what to do. Every alert must link to a runbook with diagnosis steps.
108
109## Quality Checklist
110
111```
112[ ] kube-prometheus-stack or equivalent deployed and healthy
113[ ] ServiceMonitors auto-discover all application metrics endpoints
114[ ] SLOs defined for every user-facing service
115[ ] Burn-rate alerts configured (critical, warning, ticket)
116[ ] Recording rules pre-compute expensive queries
117[ ] Grafana dashboards use RED method for services, USE for infrastructure
118[ ] Alertmanager routes to correct channels (PagerDuty/Slack/OpsGenie)
119[ ] Alert grouping and inhibition rules prevent notification storms
120[ ] Every alert has a linked runbook
121[ ] Metrics retention configured (15d local, long-term in Mimir/Thanos)
122[ ] Dashboard provisioned as code (JSON/YAML in Git)
123[ ] Error budget dashboard visible to engineering and product
124```