OnCall Alert-Group Triage
OnCall alert groups are post-routing aggregations of alert deliveries — they answer "what is paging right now?". Grafana alert rules (under gcx alert rules) answer "why is rule X evaluating to firing?". This skill covers the OnCall side and pivots out when needed.
Core Principles
- Use gcx — do not call OnCall APIs directly (no curl, no HTTP libraries).
- The CLI emits structured diagnostics on stderr (
hint:,note:,warn:; JSONL withclassin agent mode). Pass them through. - The default
alert-groups listfilter excludes resolved + child groups. Surface--allonly when the user asks about history. - Action verbs in agent mode require
--forcewhen matched count > 1. Never pre-fill--forcewithout user confirmation.
Prerequisites
gcx configured with an active context that targets a Grafana stack with OnCall enabled. If not, use the setup-gcx skill first.
Triage Workflow
Step 1: List Active Alert Groups
gcx irm oncall alert-groups list
Default filter: state in {firing, acknowledged, silenced} and is_root=true (matches the OnCall UI). See --help for the full flag set; common narrowers are --state, --team <PK>, --integration <PK>, --escalation-chain <ID>, --mine, --max-age 24h. Team / integration / chain filters take IDs, not names — resolve first:
gcx irm oncall teams list -o json | jq -r '.[] | "\(.metadata.name): \(.spec.name)"' | grep -i "my team"
gcx irm oncall escalation-chains list
To attribute alert load to a rotation, filter by escalation chain, not integration: one integration routes through several chains to several schedules, and the alert group record carries no chain field of its own. Chain-filtered queries over the same integration return disjoint subsets.
Escape hatches: --all (drops both defaults — returns resolved + child groups), --include-child-groups, --state resolved.
For a historical window use --from / --to (RFC3339, unix timestamp, or now-30d); --max-age only anchors to now, and the two cannot be combined. --resolved-from / --resolved-to bound resolved_at instead. --acknowledged-by <user-id> / --resolved-by <user-id> attribute handling to a person.
Default table: ID TITLE SEVERITY STATE TEAM SUBJECT AGE. -o wide adds RULE (Grafana rule URL) and ALERTS (group-wide count). Use -o wide when the user needs the rule pivot.
There is no --title / substring filter. To act on "all the kafka ones", fetch JSON, filter with jq, then loop the IDs through the action verb (or — if the kafka alerts all share an integration or team — bulk-by-filter with --integration <PK> / --team <PK> instead):
gcx irm oncall alert-groups list -o json | \
jq -r '.items[] | select(.status.title | test("kafka"; "i")) | .metadata.name'
Step 2: Drill Into a Group
gcx irm oncall alert-groups get <id>
One round trip already populates the rich status.links.* block — typical triage does NOT need list-alerts first.
Key paths under status:
title,severity,state(firing|acknowledged|resolved|silenced),summary,runbookURLsubject.labels— canonical commonLabels (whatever the rule grouped by; canonical only onget, best-effort onlist)timestamps.{started,acknowledged,resolved,silenced}links.alert.rule.{uid,url}— Grafana alert rule pivot (most important)links.alert.instance.{id,silenceURL}— Alertmanager fingerprintlinks.dashboard.{uid,url,panel.id,panel.url}links.slo.{uid,name}alertsCount
Also: spec.permalinks.web (OnCall UI), spec.team.{id,name}, spec.integration.{id,name,type}.
--include-raw adds the unprocessed Alertmanager payload at status.raw — only when the user needs an unpromoted label or annotation.
Step 3: Per-Alert Detail (when needed)
Only when the group has multiple firing instances and the user needs per-fire detail:
gcx irm oncall alert-groups list-alerts <id>
Default collapses by label set (Alertmanager fingerprint) — repeated fires of the same labeled instance fold into one row, status.occurrences reports the re-fire count. --history opts out (every delivery becomes a row, occurrences: 1). --slim skips the per-alert fetch for counting/sorting. --include-raw exposes the full payload. --limit default 100; the CLI warns when capped.
Per-alert key fields: status.dimensions.labels (per-fire discriminators — labels that differ from the group's subject.labels), status.occurrences, status.links.* (usually constant across siblings).
Step 4: Pivot to Alert Rule / Dashboard / SLO
The status.links.* identifiers are cross-provider pivots:
| Source field | Next command |
|---|---|
status.links.alert.rule.uid |
gcx alert instances list --rule <uid> — all currently firing instances across clusters; hand off to investigate-alert for rule-side root cause |
status.links.dashboard.uid |
gcx dashboards get <uid> (metadata/panels); gcx dashboards search <keywords> (find related); gcx dashboards snapshot <uid> --since 6h (visual inspection) |
status.links.slo.uid |
gcx slo definitions status <uid> (current SLI + budget — note: budget figure may show 100% if recording rules are absent or budget went deeply negative; use gcx dashboards snapshot grafana_slo_app-<uid> --since 28d for the authoritative historical view); then gcx slo definitions get <uid> to extract datasource UID and recording-rule queries for gcx metrics query; or slo-investigate |
Step 5: Act
The same verb runs single-target (pass <id>) or bulk-by-filter (omit <id>, pass filter flags). Verbs: acknowledge, unacknowledge, silence (+ --duration seconds), unsilence, resolve, unresolve, delete. All are idempotent except delete.
# Single-target
gcx irm oncall alert-groups acknowledge <id>
gcx irm oncall alert-groups silence <id> --duration 3600
# Bulk-by-filter (filter flags mirror `list`)
gcx irm oncall alert-groups acknowledge --team <PK> --state firing --force
Result envelopes:
// Single-target
{"action":"acknowledge","target":{"alertGroupId":"<id>"},"changed":true}
// Bulk
{"action":"acknowledge","summary":{"matched":23,"succeeded":18,"skipped":5,"failed":0},"failures":[]}
changed:false / skipped++ on idempotent re-runs (not an error). matched == succeeded + skipped + failed. failures[] carries only errored targets.
Bulk in agent mode REQUIRES --force when matched count > 1. Show the user the filter set and confirm before running — do not pre-fill --force. A bulk call with no <id> AND no filter flags is blocked with a structured error containing suggestions[].
Error Handling
<id> argument or filter flag required— surface the error'ssuggestions[]to the user.- Agent mode + matched > 1 + no
--force— confirm scope with the user, then re-run with--force. - 404 on
get <id>— group may have been purged; retry the priorlistwith--all --state resolved. webhook/formatted_webhookintegrations — partialstatus.links.*and emptysubject.labelsare expected, not errors.
Tips
getinlineslast_alert.raw_request_data— prefer it overlist-alertsfor typical triage.subject.labelsis canonical onget; onlistit is HTML-scraped (good enough for the table, not for exact-label work).- The team cell renders
<name> (<id>); the ID is the copy-paste target for--team. gcx irm oncall alertsandalerts getwere removed — usealert-groups list-alerts <group-id>.
Related Skills
- investigate-alert — rule-side root cause once you have
status.links.alert.rule.uid. - slo-investigate — when
status.links.slo.uidis populated. - debug-with-grafana — broader investigation when the alert is a symptom of a service-level issue.