Incident Triage

Interpret production incident tickets and monitoring alerts, identify the affected application or business flow, retrieve and correlate evidence (logs, telemetry, source code context, configuration, deployment history, known issues), and produce an evidence-backed incident analysis: summary, probable root-cause hypotheses with confidence indicators, impacted components, recommended diagnostic/remediation steps, and routing or escalation guidance. Triggers on phrases like: production incident, incident ticket, triage this alert, root cause, RCA, correlate logs, what broke, error spike, latency spike, transaction failed, timeout, outage, sev1/P1, escalate, on-call, anomaly detection, alert threshold, known issue, postmortem. Use whenever a production issue, alert, or incident report is being analyzed — even if the user just pastes an error message, a stack trace, a screenshot description, or asks 'why are customers seeing this error?'

aws-samples 9d82773 6 files · 31.8 KB Updated

File contents

aws-samples/sample-ramp-aidlc-mod-starter-packs/tree/main/agentic-incident-response/skills/incident-triage commit 9d82773370

Frequently asked questions

npx skillmds@latest add aws-samples/incident-triage