Prometheus Grafana Triage
Treat alerting problems as one of three classes: a real platform issue, a broken scrape, or a bad rule. Verify which one it is before touching manifests.
Use when
- Grafana alerts are firing and the user wants to know why
- Prometheus scrape targets are down
- Alertmanager is noisy or seems stale
- Grafana dashboards disagree with current cluster state
- the user suspects a rule bug or stale metric logic
Do not use when
- The primary problem is a generic Kubernetes runtime incident with no monitoring angle. Use
k8s-sre-triage. - The primary problem is a failing CI job or deployment pipeline. Use the installed GitHub
github:gh-fix-ciskill for GitHub Actions, or investigate the owning pipeline natively. - The request is for dashboard design or long-term metrics architecture rather than incident triage.
Workflow
1. Check the live monitoring stack
Verify:
- which cluster actually owns the alert evaluation or scrape path
- the expected state for each cluster, spoke, or optional runtime
- Grafana health
- Prometheus health
- Alertmanager health
- whether you are querying the hub or the spoke agent
Before switching contexts or querying random clusters, identify the source of truth:
- the cluster label on the alert or target
- the Prometheus or Alertmanager instance that is currently evaluating the rule
- whether the data comes from a hub-and-spoke flow or a local Prometheus
- repo docs, runbooks, maintenance schedules, or automation state that say a target is intentionally stopped, parked, retired, or on-demand
Do not assume the dashboard reflects the current truth until Prometheus and its scrape path are confirmed healthy.
2. Classify expected state before declaring an incident
For hub-and-spoke or cost-controlled environments, classify every relevant target before interpreting missing metrics or unreachable endpoints:
live: expected to be online nowparked: intentionally shut down, scaled to zero, paused, or on-demandretired: intentionally removed from service but still visible in old dashboards, labels, or Argo CD appsunknown intent: no current evidence says whether the target should be online
Use repo-local evidence first:
- README and architecture docs
- operations runbooks
- maintenance or startup/shutdown docs
- GitOps application descriptions
- automation schedules
- recent reports only as secondary context
For a parked or retired target, do not report absent metrics, DNS failure, Argo CD Unknown, or up == 0 as a live platform incident by itself. Report it as expected offline state, stale visibility, or an expected blind spot.
In hub-and-spoke environments, before treating missing hub data as target failure, check remote write and federation flow health first. Confirm whether there is federation lag, and inspect remote write queue depth and backfill status so ingest delay is not mistaken for a real outage.
Escalate an offline target only when:
- the user asked for that target to be checked as online
- a job, migration, maintenance window, or startup runbook expects it to be online
- recent telemetry shows it was online and then dropped unexpectedly
- the shutdown/startup automation itself reports failure
Example: if an AKS spoke is documented as on-demand for cost control, then missing AKS samples and Argo CD Healthy/Unknown are expected while it is parked. The SRE finding is the OKE hub health plus the parked-spoke state, not an AKS outage.
3. Inspect active alerts
Capture:
- alert name
- state: pending vs firing
- cluster label
- severity
- summary and description
Separate:
- active alert in Alertmanager
- active rule evaluation in Prometheus
- stale visualization in Grafana
For container restart or OOMKilled alerts, verify the runtime event separately from the alert freshness:
- current pod
restartCount lastState.reasonlastState.terminated.exitCode- previous container logs when available
Bundled helpers:
scripts/alert_summary.pyfor active alerts from Prometheus or Alertmanagerscripts/prom_target_failures.pyfor current scrape failures withlastError
4. Check scrape health
When up == 0, inspect:
- target instance
- job name
- scrape URL
lastError
Confirm the endpoint directly where possible:
- wrong path often returns
404 - wrong port or bind address often returns
connection refused - missing auth returns
401or403
For Prometheus warnings containing Error on ingesting samples with different value but same timestamp, use a duplicate-sample workflow before changing scrape config:
- map the warning target IP and port back to a pod, service, endpoint, or EndpointSlice
- sample the target directly and count duplicate metric-label-timestamp keys when the endpoint is reachable
- distinguish duplicate exposition by the target from duplicate scrape paths or overlapping scrape jobs
- inspect the owning source object or GitOps manifest that creates the duplicate target data
- after a fix, search Prometheus logs with
--since-timeafter the last known matching warning; broad log output can include stale historical warnings
5. Check the rule logic
Look for common rule problems:
- using
count(metric)when the correct intent issum(metric == 1) - alerting on sticky gauges like
last_terminated_reason == 1 - thresholds that count metric series rather than objects
- rules evaluated against the wrong cluster label
Prefer a corrected query over ad hoc silencing when the rule itself is wrong.
After any PromQL or rule fix, run deterministic rule validation with promtool check rules and promtool test rules (or equivalent rule tests) before closing the fix loop.
For OOM alerts, avoid treating a sticky last_terminated_reason gauge as proof that the problem is still active. A real OOM event can coexist with a stale firing alert.
6. Decide the fix path
Pick one:
- runtime issue: hand off to
k8s-sre-triage - scrape config issue: fix service, ServiceMonitor, scrape config, port, or path
- rule issue: fix the PromQL and docs
- temporary operational noise: before adding or updating a silence, inspect alertmanager notification policy and route configuration for routing, grouping, repeat interval, and inhibition behavior to ensure the silence maps to the right route scope
- expected offline state: document the parked, retired, or on-demand state and recommend checks only for the next startup or intended-online window
7. Verify
After the change:
- target health becomes
up - corrected query returns the expected count
- firing or pending state clears as expected
- no important alert coverage was removed accidentally
- changed rules pass deterministic checks such as
promtool check rulesandpromtool test rules - expected-state classification is backed by a current repo doc, runbook, automation schedule, or explicit user instruction
- log-derived symptoms have a cutoff proof: the same log query after the last observed warning or error timestamp returns no new matches
Guidance
- Query the system that actually scrapes the target. In a hub-and-spoke design, the spoke agent often holds the real
lastError. - Do not trust one layer alone. Compare cluster reality, Prometheus target health, and alert logic.
- Do not collapse "unreachable" into "broken" until expected state is known. On-demand infrastructure often looks broken from dashboards while parked.
- If the rule source lives in a different repo than the workload, say so clearly and patch the right repo.
container_memory_working_set_bytesis supporting evidence, not definitive proof of an exact cgroup kill threshold or kill timestamp. Use it to support an OOM diagnosis, not to overstate one.- When the evidence proves a real event but not the precise trigger, say so. Distinguish “real incident” from “fully explained incident.”
DPM and usage analysis
DPM is data points per minute per series: at a 60s scrape interval a series produces 1 DPM, at 15s it produces 4. A DPM spike with flat Active Series usually means a shortened scrape interval or duplicate remote_write senders, not new series. To find what drives it:
- Grafana Cloud billing and usage dashboards break usage down per metric and per source
count_over_time(metric[1m])orrate()on suspect metrics shows samples-per-minute directly- the tsdb status endpoints (see
prometheus-cardinality-troubleshooter) show which metrics carry the most series
Hand off to prometheus-cardinality-troubleshooter when the driver is an active series explosion, and to prometheus-label-strategy when the ask is preventing it in instrumentation or scrape config.
Related specialist skills
Use these for deeper Grafana-specific work after the incident shape is clear:
- Write and fix PromQL directly; use
prometheus-cardinality-troubleshooterwhen slowness comes from cardinality. lokifor LogQL, parsers, log-pipeline behavior, and Loki-specific troubleshooting.- Configure Grafana alert rules, contact points, notification policies, silences, and SLOs directly from Grafana Alerting knowledge.
prometheus-cardinality-troubleshooterfor active series explosions, Prometheus OOM, slow queries, ingest limits, or DPM fires.prometheus-label-strategyfor preventing cardinality problems in instrumentation or scrape target labels.loki-label-analyzerfor Loki label strategy and slow log queries caused by bad labels.- For Grafana MCP setup, follow the upstream grafana/mcp-grafana README directly.
References
- Read
references/triage-patterns.mdfor common scrape and rule failure patterns.
Scripts
scripts/alert_summary.py
Summarizes active alerts from Prometheus or Alertmanager.
Usage:
python3 "${CODEX_HOME:-$HOME/.codex}/skills/prometheus-grafana-triage/scripts/alert_summary.py" --prometheus-url http://127.0.0.1:9090
python3 "${CODEX_HOME:-$HOME/.codex}/skills/prometheus-grafana-triage/scripts/alert_summary.py" --alertmanager-url http://127.0.0.1:9093
scripts/prom_target_failures.py
Lists non-healthy Prometheus targets with cluster, job, scrape URL, and lastError.
Usage:
python3 "${CODEX_HOME:-$HOME/.codex}/skills/prometheus-grafana-triage/scripts/prom_target_failures.py" --prometheus-url http://127.0.0.1:9090