Weekly Production Review
Use this skill to produce a source-grounded production review across Datadog,
incident.io, and Linear. Keep the output factual, table-first, and easy to
audit from the linked source rows.
Scope
- Default "last week" to the previous Monday through Sunday in the user's
timezone. State both local and UTC query windows in one short scope line.
- Cover all production environments unless the user narrows scope:
prod-us, prod-eu, prod-hipaa, and prod-jp.
- The review itself changes nothing outside the tracker. Do not touch incident.io
records, follow-ups, alerts, monitors, files, Slack messages, or production
systems unless the user explicitly asks after reviewing the findings.
- Linear is the exception: evidence may be commented onto issues that already
exist, labelled and marked as agent-written. A new issue has no parent, so it
is the one write that asks first — show the set, take one go-ahead, then file
them.
linear-agent-writes is the authority;
linear-bug-triage applies it to measured
findings.
- For chat-only reviews, avoid creating report artifacts or local analysis
workspaces unless a required tool workflow explicitly does so or the user asks
for a file. If incident.io analysis tooling requires a local playbook
workspace, mention it only when relevant and keep production systems
unchanged.
- Write
No measurements found when a requested signal cannot be queried or
measured.
Related Skills
- Use
datadog-query-recipes for
production Datadog query shapes and environment/site routing.
- Use
linear-bug-triage for the Linear
write-back: it comments measured evidence onto existing issues, and files new
ones once you approve the set.
- Use
incident-alert-tickets to check
each alert cluster against the per-monitor knowledge base and record newly
root-caused clusters there as a labelled description edit.
Workflow
- Confirm the review window and timezone. If the user says "last week", use the
previous calendar week, not a rolling seven-day window.
- Gather public/customer-facing incidents from incident.io. Prefer incident.io
for accepted incidents, public incident visibility, follow-ups, and incident
status. Always produce the incident.io table below, even when no rows are
found.
- Gather incident.io alert load for the same review window. Prefer
incident.io alert or escalation stats for the primary Langfuse escalation
path/team. If the user provides an incident.io pager-load dashboard URL, use
its
escalation_path parameter as a scope hint; use the dashboard date range
only when the user explicitly scopes the review to that range instead of the
default weekly window. Group by paged engineer and incident.io time-of-day
bucket so the table shows working-hours, evening, and night load. Always
produce the incident.io alert load table below.
- Gather Linear bugs from the
bug label first. Include all bug-labeled
tickets created, updated, completed, or still open with production evidence
during the window. Inspect likely production bugs with issue details and
comments when status, owner, or evidence is unclear. Always produce the
Linear bug table below.
- Gather Datadog alert/page signals for the window. Use incident.io alerts or
escalations when they represent pages; use Datadog monitor/event data when
available. Build the exhaustive alert universe by paginating until no more
results remain for the window. Cover every production environment in scope.
Group repeated firings by monitor/page title or ID, environment, service/team,
and trigger reason.
- For every Datadog alert/page cluster, perform the deep dive before writing the
final row. Do not stop at the monitor title or count. Check the monitor's
incident-alert ticket first (see
incident-alert-tickets); a
documented cause section may explain the cluster — cite the ticket in the
incident.io / Linear Link column. Inspect matching APM
spans, representative traces, related logs, error records, exception details,
failed job logs, dependency spans, queue backlog/delay context, and monitor
time windows. Put the relevant trace/span evidence and relevant logs/errors
directly in the Datadog alerts table. Do not create a separate Datadog Issue
Deep Dives table.
- Gather Datadog error log patterns for the window. Use logs with
status:error, scope to production environments, and group by the clustered
message pattern plus service/env where available. Always produce the
Datadog logs table below. Preserve exact Datadog patterns, including wildcard
tokens, instead of paraphrasing them.
- Classify each incident, bug, alert, alert-load row, and log pattern. Separate
production breakage from self-hosted, internal-only, duplicate, canceled,
expected/test, staging/dev, monitor-noise, or unknown signals.
- Cross-reference source rows on a best-effort basis. Link Datadog rows to
matching incident.io incidents or Linear bugs when evidence supports the
relationship. Link incident.io and Linear rows back to Datadog evidence when
available. If a relationship is inferential, say so in the row.
Output Contract
Return one short scope line followed by exactly these five source tables, in
this order:
- incident.io
- incident.io Alert Load
- Linear Bugs
- Datadog Alerts
- Datadog Logs
Do not add an executive summary, narrative summary, event-centric view, summary
table, source synthesis table, or separate Datadog Issue Deep Dives table. Put
counts and classifications inside the source tables.
If a table has no rows, keep the table heading and write one row or sentence
with No rows found or No measurements found plus the scoped source/query.
If a row is unclear, classify it as unclear or unknown/no measurements
instead of dropping it.
Cross-Source Linking
Keep incident.io incidents, incident.io alert load, Linear, Datadog alerts, and
Datadog logs as separate output tables. Use links inside each table to show
relationships instead of synthesizing a separate cross-source table.
Use Linear as the source of truth for deduplication across weeks and workflows.
Before reporting a bug, security finding, cost concern, or alert as new, search
Linear for matching issue keys, titles, source URLs, and comments — covering
both the bug label set and the incident-alert label set. If an
existing issue covers it, link to that issue and mark the row as already
tracked instead of reporting it again as fresh work.
Use short stable link labels:
Datadog monitor: <monitor name>
Datadog logs: <env/service/symptom>
Datadog spans: <env/route/symptom>
Datadog trace: <trace id or route>
incident.io: <INC reference>
incident.io alert load: <escalation path or team>
Linear: <issue key>
Do not write links, create follow-ups, or update external systems unless the
user explicitly asks for changes after reviewing the report.
incident.io Table
Use this table for incident.io incidents with public or customer-facing impact
in the review window. Query incident.io with incident_list scoped to the
review window and relevant team when known. Include summary, roles,
custom_fields, timestamps, durations, and escalation_urgency when the
tool supports them. If the tool cannot filter by visibility server-side, list
the scoped incidents and include only rows where visibility is public or
the evidence supports customer-facing impact.
| Incident |
Severity / Status |
Start / End / Duration |
Impact |
Linked Sources |
Follow-ups / Notes |
Column rules:
Incident: incident.io reference linked to the incident.
Severity / Status: severity and current lifecycle status.
Start / End / Duration: reported, identified, resolved, and duration when
available.
Impact: short impact statement grounded in the incident summary.
Linked Sources: Datadog alerts/pages, Datadog logs/spans, Linear issues, or
none found.
Follow-ups / Notes: follow-up count/status or none found.
incident.io Alert Load Table
Use this table for incident.io alert and pager load in the review window.
Prefer incident.io alert or escalation stats filtered to the Langfuse escalation
path or team. If the user provides a pager-load dashboard URL, parse and apply
the escalation_path[one_of] filter when available. Count alerts when the
source returns alert counts; otherwise count escalations/pages and label the
count source in Source / Notes.
incident.io time-of-day buckets are UTC:
working_hours: 09:00-18:00 Monday-Friday.
late_evening: 18:00-23:00 any day plus weekend daytime.
overnight: 23:00-09:00.
Render one row per paged engineer, sorted by total descending, and include a
final All engineers row when measurements exist. If user identity is missing,
use Unassigned / no responder. Do not collapse this table into a narrative
summary.
| Engineer |
Working Hours |
Late Evening |
Overnight |
Total Alerts / Pages |
Share |
Source / Notes |
Column rules:
Engineer: paged user or escalation target. Use the engineer's display name
when available.
Working Hours: count in the working_hours bucket.
Late Evening: count in the late_evening bucket.
Overnight: count in the overnight bucket.
Total Alerts / Pages: row total. Label pages versus alerts in
Source / Notes when the source does not expose alert counts directly.
Share: row total divided by the measured alert/page total.
Source / Notes: source query, escalation path/team filter, dashboard link,
or No measurements found.
Linear Bugs Table
Start from all Linear tickets with the bug label that were touched by the
window. Do not rely only on text searches for prod, incident, or Datadog;
those searches are useful for enrichment but are not the source universe. This
table is the Linear source inventory. A Linear bug can be classified as
non-production, duplicate, canceled, or no-action.
| Linear |
Title |
Summary |
Owner |
Status |
Touched Last Week Because |
Production Evidence |
Classification |
Counted? |
Column rules:
Linear: issue key linked to Linear, such as LFE-123.
Title: Linear issue title in a separate column.
Summary: one operational sentence based on the issue body, comments, and
evidence. Avoid fix guesses.
Owner: assignee if present; otherwise owning team if clear; otherwise
Unassigned.
Status: Linear status or state, plus completion timing when useful.
Touched Last Week Because: created, updated, completed, or
open production bug.
Production Evidence: prod env, customer impact, incident.io incident,
Datadog link, measured logs/spans/errors, or No measurements found.
Classification: use production/customer-impacting, internal-only,
self-hosted, staging/dev, duplicate/canceled/no-action, or unclear.
Counted?: yes only when the bug label and production/customer-impacting
evidence support including it in fixed/open production bug counts.
Datadog Alerts Table
Use this table for every production Datadog alert/page cluster found in the full
paginated alert/event pass, including clusters later classified as
expected/test, monitor noise, or unknown/no measurements.
| Monitor / Page Signal |
Env / Service |
Count / Window |
Why It Alerted |
Trace / Span Evidence |
Relevant Logs / Errors |
Verdict |
incident.io / Linear Link |
Column rules:
Monitor / Page Signal: monitor/page title or stable ID.
Env / Service: production env and service/team labels.
Count / Window: grouped firing/page count and relevant time window.
Why It Alerted: monitor threshold, trigger condition, route, queue, status
code, latency, or backlog signal.
Trace / Span Evidence: representative trace/span links, error counts,
latency, status codes, dependency spans, or No measurements found.
Relevant Logs / Errors: explicit exception class/message, exact or
normalized log message, failed job IDs when visible, DB/downstream errors,
retry exhaustion, validation failures, or No measurements found.
Verdict: use customer incident, confirmed bug, infra/dependency,
expected/test, monitor noise, or unknown/no measurements.
incident.io / Linear Link: matching incident.io reference, Linear issue key,
explicit disposition, or none found.
For API route and queue consumer errors, start from APM spans matching:
operation_name:(http.server OR bullmq.consumer) status:error env:<prod-env>
Then narrow by service, resource_name, route, queue, consumer, status code,
error type, or monitor time window. For failed trace samples, explicitly query
related Datadog logs and error records using the trace ID, span ID, service,
resource name, environment, and same time window. If no related logs or error
records are found, write No measurements found.
Before finalizing, compare the final Datadog alerts table against the full
paginated alert/event sweep. Confirm every production monitor title seen during
the window appears in the table or is explicitly excluded as non-prod.
Datadog Logs Table
Use this table for the most frequent production error-log patterns from the
review window. This is a broad log-health pass and is separate from the alert
cluster rows above, though rows should cross-link when possible.
Start from a Datadog Logs query shaped like:
status:error
Scope it to the review window and production environments in scope. Prefer the
Datadog pattern/clustering view using the log message field as the clustering
pattern field. Group or facet by service, env, and status when available,
and sort by count descending. Review at least the top 10 patterns overall, plus
any additional top pattern per production environment when the global top 10 is
dominated by one env or service.
| Exact Log Pattern |
Env / Service |
Count / Share |
Representative Error |
Related Signal / Link |
Disposition |
Column rules:
Exact Log Pattern: exact Datadog clustered pattern or exact raw log message.
Preserve Datadog wildcard syntax such as [wildcard]...[/wildcard]. Do not
paraphrase this column.
Env / Service: affected production envs and services.
Count / Share: count in the review window and share if available.
Representative Error: one exact short representative message, exception
class, or stack/log summary. Avoid long stack traces.
Related Signal / Link: related Datadog alert row, trace, log query,
incident.io incident, Linear issue, or none found.
Disposition: use known incident, tracked bug, needs investigation,
expected/test, monitor noise, or unknown.
If a high-volume pattern maps to a failed API route or queue consumer, ensure
the matching Datadog alert row includes the trace/log/error investigation. If a
high-volume pattern has no alert/page row, keep it in this table anyway and
mark it needs investigation or unknown based on evidence.
Output Format
Return valid Markdown only.
- Use valid Markdown syntax for headings, links, and tables.
- Include a space after list markers such as
-, *, and 1..
- Close links and parentheses correctly.
- Do not emit malformed tables, dangling backticks, or partially opened code
fences.
- Escape table pipes inside log patterns or error messages when needed.
- If a section would be fragile to format, prefer a plain paragraph over broken
Markdown.
1---2name: weekly-production-review3description: Prepare Langfuse weekly production reviews covering failures, fixes, open issues, and tracking gaps. Use for "what broke last week," production bugs, Datadog alerts or error patterns, incident.io activity, or pager load.4---5
6# Weekly Production Review
7
8Use this skill to produce a source-grounded production review across Datadog,
9incident.io, and Linear. Keep the output factual, table-first, and easy to
10audit from the linked source rows.
11
12## Scope
13
14- Default "last week" to the previous Monday through Sunday in the user's
15 timezone. State both local and UTC query windows in one short scope line.
16- Cover all production environments unless the user narrows scope:
17 `prod-us`, `prod-eu`, `prod-hipaa`, and `prod-jp`.
18- The review itself changes nothing outside the tracker. Do not touch incident.io
19 records, follow-ups, alerts, monitors, files, Slack messages, or production
20 systems unless the user explicitly asks after reviewing the findings.
21- Linear is the exception: evidence may be commented onto issues that already
22 exist, labelled and marked as agent-written. A new issue has no parent, so it
23 is the one write that asks first — show the set, take one go-ahead, then file
24 them. [`linear-agent-writes`](../linear-agent-writes/SKILL.md) is the authority;
25 [`linear-bug-triage`](../linear-bug-triage/SKILL.md) applies it to measured
26 findings.
27- For chat-only reviews, avoid creating report artifacts or local analysis
28 workspaces unless a required tool workflow explicitly does so or the user asks
29 for a file. If incident.io analysis tooling requires a local playbook
30 workspace, mention it only when relevant and keep production systems
31 unchanged.
32- Write `No measurements found` when a requested signal cannot be queried or
33 measured.
34
35## Related Skills
36
37- Use [`datadog-query-recipes`](../datadog-query-recipes/SKILL.md) for
38 production Datadog query shapes and environment/site routing.
39- Use [`linear-bug-triage`](../linear-bug-triage/SKILL.md) for the Linear
40 write-back: it comments measured evidence onto existing issues, and files new
41 ones once you approve the set.
42- Use [`incident-alert-tickets`](../incident-alert-tickets/SKILL.md) to check
43 each alert cluster against the per-monitor knowledge base and record newly
44 root-caused clusters there as a labelled description edit.
45
46## Workflow
47
481. Confirm the review window and timezone. If the user says "last week", use the
49 previous calendar week, not a rolling seven-day window.
502. Gather public/customer-facing incidents from incident.io. Prefer incident.io
51 for accepted incidents, public incident visibility, follow-ups, and incident
52 status. Always produce the incident.io table below, even when no rows are
53 found.
543. Gather incident.io alert load for the same review window. Prefer
55 incident.io alert or escalation stats for the primary Langfuse escalation
56 path/team. If the user provides an incident.io pager-load dashboard URL, use
57 its `escalation_path` parameter as a scope hint; use the dashboard date range
58 only when the user explicitly scopes the review to that range instead of the
59 default weekly window. Group by paged engineer and incident.io time-of-day
60 bucket so the table shows working-hours, evening, and night load. Always
61 produce the incident.io alert load table below.
624. Gather Linear bugs from the `bug` label first. Include all `bug`-labeled
63 tickets created, updated, completed, or still open with production evidence
64 during the window. Inspect likely production bugs with issue details and
65 comments when status, owner, or evidence is unclear. Always produce the
66 Linear bug table below.
675. Gather Datadog alert/page signals for the window. Use incident.io alerts or
68 escalations when they represent pages; use Datadog monitor/event data when
69 available. Build the exhaustive alert universe by paginating until no more
70 results remain for the window. Cover every production environment in scope.
71 Group repeated firings by monitor/page title or ID, environment, service/team,
72 and trigger reason.
736. For every Datadog alert/page cluster, perform the deep dive before writing the
74 final row. Do not stop at the monitor title or count. Check the monitor's
75 `incident-alert` ticket first (see
76 [`incident-alert-tickets`](../incident-alert-tickets/SKILL.md)); a
77 documented cause section may explain the cluster — cite the ticket in the
78 `incident.io / Linear Link` column. Inspect matching APM
79 spans, representative traces, related logs, error records, exception details,
80 failed job logs, dependency spans, queue backlog/delay context, and monitor
81 time windows. Put the relevant trace/span evidence and relevant logs/errors
82 directly in the Datadog alerts table. Do not create a separate Datadog Issue
83 Deep Dives table.
847. Gather Datadog error log patterns for the window. Use logs with
85 `status:error`, scope to production environments, and group by the clustered
86 `message` pattern plus service/env where available. Always produce the
87 Datadog logs table below. Preserve exact Datadog patterns, including wildcard
88 tokens, instead of paraphrasing them.
898. Classify each incident, bug, alert, alert-load row, and log pattern. Separate
90 production breakage from self-hosted, internal-only, duplicate, canceled,
91 expected/test, staging/dev, monitor-noise, or unknown signals.
929. Cross-reference source rows on a best-effort basis. Link Datadog rows to
93 matching incident.io incidents or Linear bugs when evidence supports the
94 relationship. Link incident.io and Linear rows back to Datadog evidence when
95 available. If a relationship is inferential, say so in the row.
96
97## Output Contract
98
99Return one short scope line followed by exactly these five source tables, in
100this order:
101
1021. incident.io
1032. incident.io Alert Load
1043. Linear Bugs
1054. Datadog Alerts
1065. Datadog Logs
107
108Do not add an executive summary, narrative summary, event-centric view, summary
109table, source synthesis table, or separate Datadog Issue Deep Dives table. Put
110counts and classifications inside the source tables.
111
112If a table has no rows, keep the table heading and write one row or sentence
113with `No rows found` or `No measurements found` plus the scoped source/query.
114If a row is unclear, classify it as `unclear` or `unknown/no measurements`
115instead of dropping it.
116
117## Cross-Source Linking
118
119Keep incident.io incidents, incident.io alert load, Linear, Datadog alerts, and
120Datadog logs as separate output tables. Use links inside each table to show
121relationships instead of synthesizing a separate cross-source table.
122
123Use Linear as the source of truth for deduplication across weeks and workflows.
124Before reporting a bug, security finding, cost concern, or alert as new, search
125Linear for matching issue keys, titles, source URLs, and comments — covering
126both the `bug` label set and the `incident-alert` label set. If an
127existing issue covers it, link to that issue and mark the row as already
128tracked instead of reporting it again as fresh work.
129
130Use short stable link labels:
131
132- `Datadog monitor: <monitor name>`
133- `Datadog logs: <env/service/symptom>`
134- `Datadog spans: <env/route/symptom>`
135- `Datadog trace: <trace id or route>`
136- `incident.io: <INC reference>`
137- `incident.io alert load: <escalation path or team>`
138- `Linear: <issue key>`
139
140Do not write links, create follow-ups, or update external systems unless the
141user explicitly asks for changes after reviewing the report.
142
143## incident.io Table
144
145Use this table for incident.io incidents with public or customer-facing impact
146in the review window. Query incident.io with `incident_list` scoped to the
147review window and relevant team when known. Include `summary`, `roles`,
148`custom_fields`, `timestamps`, `durations`, and `escalation_urgency` when the
149tool supports them. If the tool cannot filter by visibility server-side, list
150the scoped incidents and include only rows where `visibility` is `public` or
151the evidence supports customer-facing impact.
152
153| Incident | Severity / Status | Start / End / Duration | Impact | Linked Sources | Follow-ups / Notes |
154| --- | --- | --- | --- | --- | --- |
155
156Column rules:
157
158- `Incident`: incident.io reference linked to the incident.
159- `Severity / Status`: severity and current lifecycle status.
160- `Start / End / Duration`: reported, identified, resolved, and duration when
161 available.
162- `Impact`: short impact statement grounded in the incident summary.
163- `Linked Sources`: Datadog alerts/pages, Datadog logs/spans, Linear issues, or
164 `none found`.
165- `Follow-ups / Notes`: follow-up count/status or `none found`.
166
167## incident.io Alert Load Table
168
169Use this table for incident.io alert and pager load in the review window.
170Prefer incident.io alert or escalation stats filtered to the Langfuse escalation
171path or team. If the user provides a pager-load dashboard URL, parse and apply
172the `escalation_path[one_of]` filter when available. Count alerts when the
173source returns alert counts; otherwise count escalations/pages and label the
174count source in `Source / Notes`.
175
176incident.io time-of-day buckets are UTC:
177
178- `working_hours`: 09:00-18:00 Monday-Friday.
179- `late_evening`: 18:00-23:00 any day plus weekend daytime.
180- `overnight`: 23:00-09:00.
181
182Render one row per paged engineer, sorted by total descending, and include a
183final `All engineers` row when measurements exist. If user identity is missing,
184use `Unassigned / no responder`. Do not collapse this table into a narrative
185summary.
186
187| Engineer | Working Hours | Late Evening | Overnight | Total Alerts / Pages | Share | Source / Notes |
188| --- | ---: | ---: | ---: | ---: | ---: | --- |
189
190Column rules:
191
192- `Engineer`: paged user or escalation target. Use the engineer's display name
193 when available.
194- `Working Hours`: count in the `working_hours` bucket.
195- `Late Evening`: count in the `late_evening` bucket.
196- `Overnight`: count in the `overnight` bucket.
197- `Total Alerts / Pages`: row total. Label pages versus alerts in
198 `Source / Notes` when the source does not expose alert counts directly.
199- `Share`: row total divided by the measured alert/page total.
200- `Source / Notes`: source query, escalation path/team filter, dashboard link,
201 or `No measurements found`.
202
203## Linear Bugs Table
204
205Start from all Linear tickets with the `bug` label that were touched by the
206window. Do not rely only on text searches for `prod`, `incident`, or `Datadog`;
207those searches are useful for enrichment but are not the source universe. This
208table is the Linear source inventory. A Linear bug can be classified as
209non-production, duplicate, canceled, or no-action.
210
211| Linear | Title | Summary | Owner | Status | Touched Last Week Because | Production Evidence | Classification | Counted? |
212| --- | --- | --- | --- | --- | --- | --- | --- | --- |
213
214Column rules:
215
216- `Linear`: issue key linked to Linear, such as `LFE-123`.
217- `Title`: Linear issue title in a separate column.
218- `Summary`: one operational sentence based on the issue body, comments, and
219 evidence. Avoid fix guesses.
220- `Owner`: assignee if present; otherwise owning team if clear; otherwise
221 `Unassigned`.
222- `Status`: Linear status or state, plus completion timing when useful.
223- `Touched Last Week Because`: `created`, `updated`, `completed`, or
224 `open production bug`.
225- `Production Evidence`: prod env, customer impact, incident.io incident,
226 Datadog link, measured logs/spans/errors, or `No measurements found`.
227- `Classification`: use `production/customer-impacting`, `internal-only`,
228 `self-hosted`, `staging/dev`, `duplicate/canceled/no-action`, or `unclear`.
229- `Counted?`: `yes` only when the bug label and production/customer-impacting
230 evidence support including it in fixed/open production bug counts.
231
232## Datadog Alerts Table
233
234Use this table for every production Datadog alert/page cluster found in the full
235paginated alert/event pass, including clusters later classified as
236`expected/test`, `monitor noise`, or `unknown/no measurements`.
237
238| Monitor / Page Signal | Env / Service | Count / Window | Why It Alerted | Trace / Span Evidence | Relevant Logs / Errors | Verdict | incident.io / Linear Link |
239| --- | --- | ---: | --- | --- | --- | --- | --- |
240
241Column rules:
242
243- `Monitor / Page Signal`: monitor/page title or stable ID.
244- `Env / Service`: production env and service/team labels.
245- `Count / Window`: grouped firing/page count and relevant time window.
246- `Why It Alerted`: monitor threshold, trigger condition, route, queue, status
247 code, latency, or backlog signal.
248- `Trace / Span Evidence`: representative trace/span links, error counts,
249 latency, status codes, dependency spans, or `No measurements found`.
250- `Relevant Logs / Errors`: explicit exception class/message, exact or
251 normalized log message, failed job IDs when visible, DB/downstream errors,
252 retry exhaustion, validation failures, or `No measurements found`.
253- `Verdict`: use `customer incident`, `confirmed bug`, `infra/dependency`,
254 `expected/test`, `monitor noise`, or `unknown/no measurements`.
255- `incident.io / Linear Link`: matching incident.io reference, Linear issue key,
256 explicit disposition, or `none found`.
257
258For API route and queue consumer errors, start from APM spans matching:
259
260```text
261operation_name:(http.server OR bullmq.consumer) status:error env:<prod-env>
262```
263
264Then narrow by `service`, `resource_name`, route, queue, consumer, status code,
265error type, or monitor time window. For failed trace samples, explicitly query
266related Datadog logs and error records using the trace ID, span ID, service,
267resource name, environment, and same time window. If no related logs or error
268records are found, write `No measurements found`.
269
270Before finalizing, compare the final Datadog alerts table against the full
271paginated alert/event sweep. Confirm every production monitor title seen during
272the window appears in the table or is explicitly excluded as non-prod.
273
274## Datadog Logs Table
275
276Use this table for the most frequent production error-log patterns from the
277review window. This is a broad log-health pass and is separate from the alert
278cluster rows above, though rows should cross-link when possible.
279
280Start from a Datadog Logs query shaped like:
281
282```text
283status:error
284```
285
286Scope it to the review window and production environments in scope. Prefer the
287Datadog pattern/clustering view using the log `message` field as the clustering
288pattern field. Group or facet by `service`, `env`, and `status` when available,
289and sort by count descending. Review at least the top 10 patterns overall, plus
290any additional top pattern per production environment when the global top 10 is
291dominated by one env or service.
292
293| Exact Log Pattern | Env / Service | Count / Share | Representative Error | Related Signal / Link | Disposition |
294| --- | --- | ---: | --- | --- | --- |
295
296Column rules:
297
298- `Exact Log Pattern`: exact Datadog clustered pattern or exact raw log message.
299 Preserve Datadog wildcard syntax such as `[wildcard]...[/wildcard]`. Do not
300 paraphrase this column.
301- `Env / Service`: affected production envs and services.
302- `Count / Share`: count in the review window and share if available.
303- `Representative Error`: one exact short representative message, exception
304 class, or stack/log summary. Avoid long stack traces.
305- `Related Signal / Link`: related Datadog alert row, trace, log query,
306 incident.io incident, Linear issue, or `none found`.
307- `Disposition`: use `known incident`, `tracked bug`, `needs investigation`,
308 `expected/test`, `monitor noise`, or `unknown`.
309
310If a high-volume pattern maps to a failed API route or queue consumer, ensure
311the matching Datadog alert row includes the trace/log/error investigation. If a
312high-volume pattern has no alert/page row, keep it in this table anyway and
313mark it `needs investigation` or `unknown` based on evidence.
314
315## Output Format
316
317Return valid Markdown only.
318
319- Use valid Markdown syntax for headings, links, and tables.
320- Include a space after list markers such as `-`, `*`, and `1.`.
321- Close links and parentheses correctly.
322- Do not emit malformed tables, dangling backticks, or partially opened code
323 fences.
324- Escape table pipes inside log patterns or error messages when needed.
325- If a section would be fragile to format, prefer a plain paragraph over broken
326 Markdown.