Incident Alert Tickets
Incident-alert tickets turn on-call debugging into searchable knowledge: one
Linear ticket per production alert/monitor, one dated section per distinct
root cause. This skill owns lookup, comparison, and the human-gated
write-back; calling skills own the investigation itself. The team SOP is the
Linear document titled "Incident Alert Tickets".
When to Apply
Apply whenever the task is anchored to an alert identity:
- a Datadog monitor ID or title,
- an incident.io alert or INC reference,
- an on-call page ("we got paged for X").
Do not try to detect "incident mode" — the presence of a named alert is the
condition, because the monitor is the ticket key. When an alert identity is
present, the lookup is mandatory; recording is offered after the investigation
and gated on human approval. A customer report or code question with no alert
identity skips this skill.
When multiple alerts fire together (a cascade), run lookup, compare, and
classify for each alert identity — every monitor has its own ticket. If
one root cause explains several alerts, write the full cause section on the
monitor closest to the cause and propose a short dated section on the other
monitors' tickets that links to it.
Ticket Contract
One ticket per monitor (per env when monitors are per-env), titled
[ENV] <Monitor title>, in the Engineering (LFE) team, carrying the
incident-alert label. The label set is the knowledge base.
Regional twins of one monitor (same metric and threshold per env) may share
a single ticket titled [ENV1/ENV2] <Monitor title> when the causes are
region-independent; list each env's monitor ID in the alert header.
The description opens with an alert header: monitor ID, trigger condition,
and how it surfaces (incident.io urgency, auto-resolve behavior).
One dated section per distinct root cause, separated by ---:
## YYYY-MM-DD — <short cause name>
**Recognize it:** <signals that identify this cause: log patterns, span
filters, metric shapes, affected routes>
**How urgent?** <impact, auto-recovery behavior, escalation threshold>
**Fix:** <positive actions only — every "do not X" needs a working
alternative; verified levers, not speculation>
Cause sections are append-only: never rewrite or delete an existing section;
new knowledge gets a new dated block.
The description ends with a ## Your cause is not listed? trailer: it
records firings that were never root-caused and tells the next engineer to
insert new dated sections above it, in the same format.
Keep each cause section to roughly one screen.
A distinct problem discovered during the investigation that is not a cause
of this alert gets its own ticket (bug or incident-alert), cross-linked — do
not mix it into this ticket's cause sections.
Lookup
- List Linear issues carrying the
incident-alert label.
- Match on monitor ID first (tickets carry it in the alert header), then on
monitor title and env.
- Read the matched ticket's cause sections and comments.
Compare and Classify
Compare the current evidence against each cause section's "Recognize it"
signals and classify:
- Known cause — a section matches. Cite the ticket and section in the
analysis; its "Fix" is the starting recommendation. This may end the
investigation before any Datadog sweep.
- New cause on existing ticket — the monitor has a ticket but no section
matches the evidence. Propose appending a dated section.
- No ticket — no ticket matches the monitor. Propose creating one, and file
it once that is approved.
Treat a partial match — some "Recognize it" signals fit, others do not — as a
new cause, never as a known one: do not recommend a documented "Fix"
whose recognition signals only partially match. Name the near-miss section in
the proposed ticket so the human can judge the overlap.
Write-Back
Appending a cause section to a ticket that already exists is a description edit;
creating a monitor with no ticket is a parentless create. Both need a yes —
see linear-agent-writes and read it before
your first write.
- Append: show it, then do it. Prepare the new
----separated dated block
(insert after the existing cause sections, above the Your cause is not listed? trailer; leave everything else untouched). Mark the block as
agent-written in its own text. Put it in the go-ahead table; once approved,
append it and label the ticket AI edited. Never reflow or rewrite the
human-written prose around it.
- Create: show it, then file it. Prepare the issue — title
[ENV] <Monitor title>, the incident-alert label, description = alert header,
the first dated cause section, and the Your cause is not listed? trailer —
show it for a go-ahead, and once you have one, create it yourself and label it
AI created.
One go-ahead covers the whole run — appends and creates together. Asking
per row is worse than the pasting this replaced.
Report what you did either way:
| ID |
Alert / Monitor |
Classification |
Action |
Content |
Action: appended to <key> (AI edited), awaiting your go-ahead,
filed <key> (AI created), or none (known cause).
Content: the dated section, or the full ticket body exactly as it will be
filed — that text is what the go-ahead is given against.
If Linear is unreachable in this environment, say so and return every row as text
ready to paste rather than skipping the write-back silently.
Division of Labor
linear-bug-triage owns bug deduplication
and creation from measured evidence. Incident-alert tickets are per-monitor
runbook knowledge, not defect reports.
- An alert whose root cause is a code bug gets both: the cause section
documents recognition and mitigation, and links the bug ticket that tracks
the durable fix.
1---2name: incident-alert-tickets3description: Read and record root causes in the Linear `incident-alert` knowledge base. Use before and after investigating a named Datadog monitor, incident.io alert or incident, or on-call page to find or record root causes.4---56# Incident Alert Tickets78Incident-alert tickets turn on-call debugging into searchable knowledge: one9Linear ticket per production alert/monitor, one dated section per distinct10root cause. This skill owns lookup, comparison, and the human-gated11write-back; calling skills own the investigation itself. The team SOP is the12Linear document titled "Incident Alert Tickets".1314## When to Apply1516Apply whenever the task is anchored to an **alert identity**:1718- a Datadog monitor ID or title,19- an incident.io alert or INC reference,20- an on-call page ("we got paged for X").2122Do not try to detect "incident mode" — the presence of a named alert is the23condition, because the monitor is the ticket key. When an alert identity is24present, the lookup is mandatory; recording is offered after the investigation25and gated on human approval. A customer report or code question with no alert26identity skips this skill.2728When multiple alerts fire together (a cascade), run lookup, compare, and29classify for **each** alert identity — every monitor has its own ticket. If30one root cause explains several alerts, write the full cause section on the31monitor closest to the cause and propose a short dated section on the other32monitors' tickets that links to it.3334## Ticket Contract3536- One ticket per monitor (per env when monitors are per-env), titled37 `[ENV] <Monitor title>`, in the Engineering (LFE) team, carrying the38 `incident-alert` label. The label set is the knowledge base.39- Regional twins of one monitor (same metric and threshold per env) may share40 a single ticket titled `[ENV1/ENV2] <Monitor title>` when the causes are41 region-independent; list each env's monitor ID in the alert header.42- The description opens with an alert header: monitor ID, trigger condition,43 and how it surfaces (incident.io urgency, auto-resolve behavior).44- One dated section per distinct root cause, separated by `---`:4546 ```markdown47 ## YYYY-MM-DD — <short cause name>4849 **Recognize it:** <signals that identify this cause: log patterns, span50 filters, metric shapes, affected routes>5152 **How urgent?** <impact, auto-recovery behavior, escalation threshold>5354 **Fix:** <positive actions only — every "do not X" needs a working55 alternative; verified levers, not speculation>56 ```5758- Cause sections are append-only: never rewrite or delete an existing section;59 new knowledge gets a new dated block.60- The description ends with a `## Your cause is not listed?` trailer: it61 records firings that were never root-caused and tells the next engineer to62 insert new dated sections above it, in the same format.63- Keep each cause section to roughly one screen.64- A distinct problem discovered during the investigation that is *not* a cause65 of this alert gets its own ticket (bug or incident-alert), cross-linked — do66 not mix it into this ticket's cause sections.6768## Lookup69701. List Linear issues carrying the `incident-alert` label.712. Match on monitor ID first (tickets carry it in the alert header), then on72 monitor title and env.733. Read the matched ticket's cause sections and comments.7475## Compare and Classify7677Compare the current evidence against each cause section's "Recognize it"78signals and classify:7980- **Known cause** — a section matches. Cite the ticket and section in the81 analysis; its "Fix" is the starting recommendation. This may end the82 investigation before any Datadog sweep.83- **New cause on existing ticket** — the monitor has a ticket but no section84 matches the evidence. Propose appending a dated section.85- **No ticket** — no ticket matches the monitor. Propose creating one, and file86 it once that is approved.8788Treat a partial match — some "Recognize it" signals fit, others do not — as a89**new cause**, never as a known one: do not recommend a documented "Fix"90whose recognition signals only partially match. Name the near-miss section in91the proposed ticket so the human can judge the overlap.9293## Write-Back9495Appending a cause section to a ticket that already exists is a description edit;96creating a monitor with no ticket is a parentless create. Both need a yes —97see [`linear-agent-writes`](../linear-agent-writes/SKILL.md) and read it before98your first write.99100- **Append: show it, then do it.** Prepare the new `---`-separated dated block101 (insert after the existing cause sections, above the `Your cause is not102 listed?` trailer; leave everything else untouched). Mark the block as103 agent-written in its own text. Put it in the go-ahead table; once approved,104 append it and label the ticket `AI edited`. Never reflow or rewrite the105 human-written prose around it.106- **Create: show it, then file it.** Prepare the issue — title107 `[ENV] <Monitor title>`, the `incident-alert` label, description = alert header,108 the first dated cause section, and the `Your cause is not listed?` trailer —109 show it for a go-ahead, and once you have one, create it yourself and label it110 `AI created`.111112**One go-ahead covers the whole run** — appends and creates together. Asking113per row is worse than the pasting this replaced.114115Report what you did either way:116117| ID | Alert / Monitor | Classification | Action | Content |118| --- | --- | --- | --- | --- |119120- `Action`: `appended to <key> (AI edited)`, `awaiting your go-ahead`,121 `filed <key> (AI created)`, or `none (known cause)`.122- `Content`: the dated section, or the full ticket body exactly as it will be123 filed — that text is what the go-ahead is given against.124125If Linear is unreachable in this environment, say so and return every row as text126ready to paste rather than skipping the write-back silently.127128## Division of Labor129130- [`linear-bug-triage`](../linear-bug-triage/SKILL.md) owns bug deduplication131 and creation from measured evidence. Incident-alert tickets are per-monitor132 runbook knowledge, not defect reports.133- An alert whose root cause is a code bug gets both: the cause section134 documents recognition and mitigation, and links the bug ticket that tracks135 the durable fix.