1---2name: prometheus3description: Prometheus monitoring patterns, cardinality management, alerting best practices, and PromQL traps.4---56## Cardinality Explosions78- Every unique label combination creates a new time series — `user_id` as label kills Prometheus9- Avoid high-cardinality labels: user IDs, email addresses, request IDs, timestamps, UUIDs10- Check cardinality: `prometheus_tsdb_head_series` metric — above 1M series needs attention11- Use histograms for latency, not per-request labels — buckets are fixed cardinality12- Relabeling can drop dangerous labels before ingestion: `labeldrop` in scrape config1314## Histogram vs Summary1516- Histograms: use for SLOs, aggregatable across instances, buckets defined upfront17- Summaries: use when you need exact percentiles, cannot aggregate across instances18- Histogram bucket boundaries must be defined before data arrives — wrong buckets = wrong percentiles19- Default buckets (.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5, 10) assume HTTP latency — adjust for your use case2021## Rate and Increase2223- `rate()` requires range selector at least 4x scrape interval — `rate(metric[1m])` with 30s scrape misses data24- `rate()` is per-second, `increase()` is total over range — don't confuse them25- Counter resets on restart — `rate()` handles this, raw delta doesn't26- `irate()` uses only last two samples — too spiky for alerting, use `rate()` for alerts2728## Alerting Mistakes2930- Alert on symptoms, not causes — "high latency" not "high CPU"31- `for` clause prevents flapping: `for: 5m` means condition must hold 5 minutes before firing32- Missing `for` clause = fires immediately on first match = noisy33- Alerts need `runbook_url` label — on-call needs to know what to do, not just that something's wrong34- Test alerts with `promtool check rules` — syntax errors discovered at 3am are bad3536## PromQL Traps3738- `and` is intersection by labels, not boolean AND — results must have matching label sets39- `or` fills in missing series, doesn't do boolean OR on values40- `{}` without metric name is expensive — scans all metrics41- `offset` goes back in time: `metric offset 1h` is value from 1 hour ago42- Comparison operators filter series: `http_requests > 100` drops series below 100, doesn't return boolean4344## Scrape Configuration4546- `honor_labels: true` trusts source labels — use only when source is authoritative (e.g., Pushgateway)47- `scrape_timeout` must be less than `scrape_interval` — otherwise overlapping scrapes48- Static configs don't reload without restart — use file_sd or service discovery for dynamic targets49- TLS verification disabled (`insecure_skip_verify`) should be temporary, never permanent5051## Pushgateway Pitfalls5253- Pushgateway is for batch jobs, not services — services should expose /metrics54- Metrics persist until deleted — stale metrics from dead jobs confuse dashboards55- Add job and instance labels to distinguish sources — default grouping hides failures56- Delete metrics when job completes: `curl -X DELETE http://pushgateway/metrics/job/myjob`5758## Recording Rules5960- Pre-compute expensive queries: `record: job:request_duration_seconds:rate5m`61- Naming convention: `level:metric:operations` — helps identify what rules produce62- Recording rules update every evaluation interval — not instant, plan for slight delay63- Reduce cardinality with recording rules: aggregate away labels you don't need for alerting6465## Federation and Remote Write6667- Federation for pulling from other Prometheus — use sparingly, adds latency68- Remote write for long-term storage — Prometheus local storage is not durable69- Remote write can buffer during outages — but buffer is finite, data loss on extended outages70- Prometheus is not highly available by default — run two instances scraping same targets7172## Common Operational Issues7374- TSDB corruption on unclean shutdown — use `--storage.tsdb.wal-compression` and monitor disk space75- Memory grows with series count — each series costs ~3KB RAM76- Compaction pauses during high load — leave 40% disk headroom77- Scrape targets stuck "Unknown" — check network, firewall, target actually exposing /metrics7879## Label Best Practices8081- Use labels for dimensions you'll filter/aggregate by — environment, service, instance82- Keep label values low-cardinality — tens or hundreds, not thousands83- Consistent naming: `snake_case`, prefix with domain: `http_requests_total`, `node_cpu_seconds_total`84- `le` label is reserved for histogram buckets — don't use for other purposes