triage-agent-events
Private developer skill for triaging crashes, panics, and
fatals across the Netdata fleet. Reads the agent-events
systemd-journal namespace via the Netdata systemd-journal
Function (Cloud-proxied or direct-agent transport) and ships
scripts that bake in index-friendly query patterns.
Why this skill exists
40k-200k status events arrive on the ingestion server every
day on stable releases. The total fleet is 1.5M agents, so
the dataset is large and noisy (many unupdated agents report
crashes that have been fixed). Naive "grep all" queries are
slow and wasteful. This skill teaches the maintainer (and any
AI assistant helping them) how to slice the dataset
efficiently and how to interpret what comes back.
Workflow
+-------------------------+ +---------------------+
| get-events.sh | --> | <timestamp>.json |
| (cloud or agent API) | | in .local/audits |
+-------------------------+ +---------------------+
|
v
+------------------------+
| analyze-events.sh |
| --by signal|version| |
| function|... |
+------------------------+
|
v
+------------------------+
| cluster + read source |
| + fix the bug |
+------------------------+
The skill is a bug-investigation tool, not a generic logs
query tool. The two existing query-netdata-cloud and
query-netdata-agents skills already cover transport
mechanics; this skill EXTENDS them with the agent-events
specifics (what fields are present, what predicates are
index-friendly, what each enum value means for triage).
Key concepts (read first)
The dataset: 40k-200k status events / day on stable
releases, spread across 1.5M agents (not all restart
daily). Naive full-namespace queries with bare FTS are
slow.
Index-friendly queries (HARD RULE): use multi-value
field filters FIRST. The Netdata systemd-journal plugin
supports the syntax:
(FIELD1 in A, B, C) AND (FIELD2 in D, E, F) AND ...
Between fields = AND. Between values = OR. This is a
facet-engine feature, NOT raw journalctl. Use FTS via
query= only as a residual narrower over the structured
slice. See query-discipline.md.
Three transports (priority order):
- Cloud API -- proxied through Netdata Cloud at the
agent-events space. Primary for the team.
- Direct agent API -- against the agent-events node's
/api/v3/function?function=systemd-journal. Primary for
scripts.
- ssh to the host -- operator-only path; mentioned in
transports.md but no scripted ssh transport.
After-the-fact event model: agents POST events ONLY on
start (the previous session's exit reason). They commit
status to disk on start, stop, and at most every 10
minutes. So the meaningful query unit is "events posted in
the last 24 hours"; "the last hour" misses real crashes
that haven't restarted yet.
23h client-side dedup (src/daemon/status-file-dedup.c:11):
same agent + same event-content hash within 23h ->
suppressed at the producer. So 1 record per agent per
event-signature per day is the natural unit. Different
agents posting the same crash signature -> both arrive
(server does not dedup).
Default time + version filters: 24h time window;
latest stable + latest 2-3 nightlies for version. This
focuses triage on bugs that still matter. Wide windows
(--since '7d' or longer) are reserved for rare crashes
(1-per-few-days class) and for "when did this start /
get fixed" investigations.
AE_ field naming*: every JSON path in the producer's
status document becomes an AE_-prefixed journal field
(per log2journal --prefix 'AE_' on the ingestion server).
See AE_FIELDS.md for the verified map and enum meanings.
Table of contents
| Doc |
Purpose |
AE_FIELDS.md |
Verified field map (~80 rows) + enum meanings for triage. Indispensable. |
transports.md |
Cloud API + direct agent API call patterns; ssh footnote. |
update-cadence.md |
After-the-fact model, dedup, push timing, disk commits, query implications. |
query-discipline.md |
The multi-value filter syntax, structured-filters-first rule, anti-patterns. |
finding-crashes.md |
Recipe: signal crashes (SIGSEGV / SIGBUS / SIGFPE / SIGABRT) on stable. |
finding-fatals.md |
Recipe: deliberate fatals (OOM, disk full, asserts). |
recipes/INDEX.md |
Live catalog of recipes (find-by-function, find-by-version, find-related-to-work). |
how-tos/INDEX.md |
Live catalog: every analysis question becomes a how-to entry. |
Live how-to rule (mandatory)
If asked a concrete question about agent-events that requires
non-trivial analysis (multiple file reads, running queries,
cross-referencing with producer source) AND the answer is not
already documented in the per-domain guides above or in
recipes/, the assistant MUST author a new how-to under
how-tos/<slug>.md and add a one-line entry to
how-tos/INDEX.md BEFORE completing the task.
Scripts (in scripts/)
| Script |
Purpose |
_lib.sh |
Helpers (agentevents_* prefix). Sources query-netdata-agents/scripts/_lib.sh. Token-safe; ships a no-leak self-test. |
get-events.sh |
Fetch events of interest. Index-friendly defaults. JSON output to .local/audits/query-agent-events/. |
analyze-events.sh |
Group-by stats over a downloaded dump (signal, version, fatal_function, architecture, etc.). |
redact-events.sh |
Opt-in redaction (machine_guid / claim_id / host_id / ephemeral_id -> placeholders). For sharing only. |
Path discipline
This skill follows
<repo>/.agents/sensitive-data-discipline.md:
- Repo files: repo-relative (
<repo>/src/...).
- Sibling Netdata-org repos:
${NETDATA_REPOS_DIR}/<repo>/....
- agent-events host / namespace / machine GUID / node ID:
ALWAYS via env keys. Never literal values in any committed
file.
- Producer ingest URL: NEVER quoted literally. Reference only
as
src/daemon/status-file.c:988.
- Fetched event payloads land under
<repo>/.local/audits/query-agent-events/<timestamp>.json
(gitignored). Do NOT paste raw event JSON into committed
artifacts.
Required env keys
| Key |
Role |
NETDATA_CLOUD_TOKEN |
Cloud REST token (long-lived). |
NETDATA_CLOUD_HOSTNAME |
Cloud REST API host. |
AGENT_EVENTS_HOSTNAME |
Dual-duty: ssh host AND direct-HTTP host of the ingestion node. Can be IP or DNS name. NOT the journalctl namespace (hardcoded agent-events); NOT the Cloud room name (also hardcoded agent-events). |
AGENT_EVENTS_MACHINE_GUID |
Agent machine GUID for direct-agent transport. |
AGENT_EVENTS_NODE_ID |
Cloud node UUID for cloud-proxy transport. |
All values live in <repo>/.env (gitignored). See
<repo>/.agents/ENV.md for setup (where each value comes
from, sample formats, common mistakes).
Related skills
query-netdata-cloud -- transport: Cloud REST API.
query-netdata-agents -- transport: direct agent REST + bearer auto-mint.
- This skill consumes both via their
_lib.sh helpers.
1---2name: triage-agent-events3description: Bug-investigation tool for the Netdata agent-events ingestion namespace -- triage crashes, panics, fatals across the fleet by downloading events of interest and clustering locally. Covers the three transports (Cloud API and direct agent API are primary; ssh is operator-only), the verified AE_* field map and enum meanings, the dedup model (23h client-side per agent and event signature), the after-the-fact event timing (POST only on agent restart), and the Netdata systemd-journal plugin multi-value filter syntax (FIELD in A, B, C) AND ... Use when investigating crashes / panics / fatals; when grepping for events touching a specific function or file or version; when looking for regressions across versions; when an agent is reported crashing in a way you want to triage. Ships scripts get-events.sh and analyze-events.sh that fetch events with index-friendly filters and compute group-by stats. Defaults to last 24 hours and to the latest stable plus latest 2-3 nightlies.4---5
6# triage-agent-events
7
8Private developer skill for triaging crashes, panics, and
9fatals across the Netdata fleet. Reads the agent-events
10systemd-journal namespace via the Netdata `systemd-journal`
11Function (Cloud-proxied or direct-agent transport) and ships
12scripts that bake in index-friendly query patterns.
13
14## Why this skill exists
15
1640k-200k status events arrive on the ingestion server every
17day on stable releases. The total fleet is 1.5M agents, so
18the dataset is large and noisy (many unupdated agents report
19crashes that have been fixed). Naive "grep all" queries are
20slow and wasteful. This skill teaches the maintainer (and any
21AI assistant helping them) how to slice the dataset
22efficiently and how to interpret what comes back.
23
24## Workflow
25
26```
27+-------------------------+ +---------------------+
28| get-events.sh | --> | <timestamp>.json |
29| (cloud or agent API) | | in .local/audits |
30+-------------------------+ +---------------------+
31 |
32 v
33 +------------------------+
34 | analyze-events.sh |
35 | --by signal|version| |
36 | function|... |
37 +------------------------+
38 |
39 v
40 +------------------------+
41 | cluster + read source |
42 | + fix the bug |
43 +------------------------+
44```
45
46The skill is a bug-investigation tool, not a generic logs
47query tool. The two existing `query-netdata-cloud` and
48`query-netdata-agents` skills already cover transport
49mechanics; this skill EXTENDS them with the agent-events
50specifics (what fields are present, what predicates are
51index-friendly, what each enum value means for triage).
52
53## Key concepts (read first)
54
551. **The dataset**: 40k-200k status events / day on stable
56 releases, spread across 1.5M agents (not all restart
57 daily). Naive full-namespace queries with bare FTS are
58 slow.
59
602. **Index-friendly queries** (HARD RULE): use multi-value
61 field filters FIRST. The Netdata `systemd-journal` plugin
62 supports the syntax:
63 ```
64 (FIELD1 in A, B, C) AND (FIELD2 in D, E, F) AND ...
65 ```
66 Between fields = AND. Between values = OR. This is a
67 facet-engine feature, NOT raw journalctl. Use FTS via
68 `query=` only as a residual narrower over the structured
69 slice. See `query-discipline.md`.
70
713. **Three transports** (priority order):
72 - **Cloud API** -- proxied through Netdata Cloud at the
73 agent-events space. Primary for the team.
74 - **Direct agent API** -- against the agent-events node's
75 `/api/v3/function?function=systemd-journal`. Primary for
76 scripts.
77 - **ssh to the host** -- operator-only path; mentioned in
78 `transports.md` but no scripted ssh transport.
79
804. **After-the-fact event model**: agents POST events ONLY on
81 start (the previous session's exit reason). They commit
82 status to disk on start, stop, and at most every 10
83 minutes. So the meaningful query unit is "events posted in
84 the last 24 hours"; "the last hour" misses real crashes
85 that haven't restarted yet.
86
875. **23h client-side dedup** (`src/daemon/status-file-dedup.c:11`):
88 same agent + same event-content hash within 23h ->
89 suppressed at the producer. So 1 record per agent per
90 event-signature per day is the natural unit. Different
91 agents posting the same crash signature -> both arrive
92 (server does not dedup).
93
946. **Default time + version filters**: 24h time window;
95 latest stable + latest 2-3 nightlies for version. This
96 focuses triage on bugs that still matter. Wide windows
97 (`--since '7d'` or longer) are reserved for rare crashes
98 (1-per-few-days class) and for "when did this start /
99 get fixed" investigations.
100
1017. **AE_* field naming**: every JSON path in the producer's
102 status document becomes an `AE_`-prefixed journal field
103 (per `log2journal --prefix 'AE_'` on the ingestion server).
104 See `AE_FIELDS.md` for the verified map and enum meanings.
105
106## Table of contents
107
108| Doc | Purpose |
109|---|---|
110| `AE_FIELDS.md` | Verified field map (~80 rows) + enum meanings for triage. Indispensable. |
111| `transports.md` | Cloud API + direct agent API call patterns; ssh footnote. |
112| `update-cadence.md` | After-the-fact model, dedup, push timing, disk commits, query implications. |
113| `query-discipline.md` | The multi-value filter syntax, structured-filters-first rule, anti-patterns. |
114| `finding-crashes.md` | Recipe: signal crashes (SIGSEGV / SIGBUS / SIGFPE / SIGABRT) on stable. |
115| `finding-fatals.md` | Recipe: deliberate fatals (OOM, disk full, asserts). |
116| `recipes/INDEX.md` | Live catalog of recipes (find-by-function, find-by-version, find-related-to-work). |
117| `how-tos/INDEX.md` | Live catalog: every analysis question becomes a how-to entry. |
118
119## Live how-to rule (mandatory)
120
121If asked a concrete question about agent-events that requires
122non-trivial analysis (multiple file reads, running queries,
123cross-referencing with producer source) AND the answer is not
124already documented in the per-domain guides above or in
125`recipes/`, the assistant MUST author a new how-to under
126`how-tos/<slug>.md` and add a one-line entry to
127`how-tos/INDEX.md` BEFORE completing the task.
128
129## Scripts (in scripts/)
130
131| Script | Purpose |
132|---|---|
133| `_lib.sh` | Helpers (`agentevents_*` prefix). Sources `query-netdata-agents/scripts/_lib.sh`. Token-safe; ships a no-leak self-test. |
134| `get-events.sh` | Fetch events of interest. Index-friendly defaults. JSON output to `.local/audits/query-agent-events/`. |
135| `analyze-events.sh` | Group-by stats over a downloaded dump (signal, version, fatal_function, architecture, etc.). |
136| `redact-events.sh` | Opt-in redaction (machine_guid / claim_id / host_id / ephemeral_id -> placeholders). For sharing only. |
137
138## Path discipline
139
140This skill follows
141`<repo>/.agents/sensitive-data-discipline.md`:
142
143- Repo files: repo-relative (`<repo>/src/...`).
144- Sibling Netdata-org repos: `${NETDATA_REPOS_DIR}/<repo>/...`.
145- agent-events host / namespace / machine GUID / node ID:
146 ALWAYS via env keys. Never literal values in any committed
147 file.
148- Producer ingest URL: NEVER quoted literally. Reference only
149 as `src/daemon/status-file.c:988`.
150- Fetched event payloads land under
151 `<repo>/.local/audits/query-agent-events/<timestamp>.json`
152 (gitignored). Do NOT paste raw event JSON into committed
153 artifacts.
154
155## Required env keys
156
157| Key | Role |
158|---|---|
159| `NETDATA_CLOUD_TOKEN` | Cloud REST token (long-lived). |
160| `NETDATA_CLOUD_HOSTNAME` | Cloud REST API host. |
161| `AGENT_EVENTS_HOSTNAME` | Dual-duty: ssh host AND direct-HTTP host of the ingestion node. Can be IP or DNS name. NOT the journalctl namespace (hardcoded `agent-events`); NOT the Cloud room name (also hardcoded `agent-events`). |
162| `AGENT_EVENTS_MACHINE_GUID` | Agent machine GUID for direct-agent transport. |
163| `AGENT_EVENTS_NODE_ID` | Cloud node UUID for cloud-proxy transport. |
164
165All values live in `<repo>/.env` (gitignored). See
166`<repo>/.agents/ENV.md` for setup (where each value comes
167from, sample formats, common mistakes).
168
169## Related skills
170
171- `query-netdata-cloud` -- transport: Cloud REST API.
172- `query-netdata-agents` -- transport: direct agent REST + bearer auto-mint.
173- This skill consumes both via their `_lib.sh` helpers.