AI Incident Postmortem Template Skill
Use When
- Produce or update AI incident postmortem from approved project evidence.
- Resolve decisions about source-attributed timeline, impact, taxonomy-based causes, actions, and publication decision.
- Prepare a reviewable handoff for Service owners, governance, and customers as approved.
Do Not Use When
- The task is primarily owned by incident-response-runbook; route there and use this skill only for its named output.
- Required project evidence or decision authority is unavailable and the requester expects a pass, release, certification, or production change.
Required Inputs
| Artefact |
Source/provider |
Required? |
Behaviour when absent |
| Project _context/, approved requirements, and relevant architecture |
Project owner and upstream phase skills |
Required |
Stop at a gap register; do not invent scope, thresholds, integrations, or owners. |
| Existing artefact, implementation, configuration, and evidence named below |
Repository, delivery team, or service owner |
Required when updating or assessing |
Mark inaccessible items not assessed; do not treat them as passed. |
| Target audience, environment, risk tolerance, and authority |
Requester and accountable owner |
Required |
Produce a read-only outline with explicit assumptions; do not mutate project or production state. |
Outputs
| Artefact |
Consumer |
Observable acceptance condition |
| AI Incident Postmortem |
Service owners, governance, and customers as approved |
The postmortem is blameless, evidence-linked, names residual risk, and assigns each corrective action an owner and due date. |
| Decision and gap register |
Reviewer and downstream phase owner |
Every assumption, rejected option, unresolved dependency, waiver, and owner is explicit. |
| Validation evidence |
Release or governance reviewer |
Checks identify command or method, date, result, evidence location, and all unassessed items. |
Evidence Produced
| Evidence |
Minimum content |
Acceptance |
| Traceability record |
Source artefact, decision, output section, owner |
No mandatory decision is source-free. |
| Quality-gate result |
Check, expected result, observed result, evidence path |
Failures and unavailable checks cannot appear as passes. |
| Review record |
Reviewer, date, disposition, open actions |
The consumer can reproduce the acceptance decision. |
Capability and Permission Boundaries
- Minimum capabilities: read and search the authorised project sources. Execution is optional and limited to non-destructive validation.
- Assessment and planning default to read-only. Create or edit the named project document only when the request explicitly authorises it. Production mutation, publishing, destructive action, spending, external communication, or certification claims require separate explicit authority.
- Treat secrets, tenant data, incident evidence, and financial records as least-privilege inputs; expose only the minimum evidence needed for review.
Degraded Mode
If files, execution, network, rendering, environment access, fonts, or current evidence are unavailable, return the narrowest useful draft plus a gap register. Label affected checks not assessed, retain the intended acceptance oracle, and state who must supply or verify the missing evidence. Never convert an unavailable check into a pass.
Decision Rules
| Choice |
Action |
Failure or risk avoided |
| Evidence is complete and authority is explicit |
Choose conclusions only from preserved incident evidence and produce the full artefact. |
Blame, hindsight narratives, or unsupported causes. |
| A required source or approval is missing |
Stop the affected branch; record the gap, owner, and unblock condition. |
Fabricated requirements or unauthorised action. |
| Evidence conflicts across sources |
Preserve both claims, identify the controlling owner, and request a recorded decision. |
Silent selection of a convenient but wrong source. |
| A check cannot run in the available environment |
Keep its oracle and mark it not assessed; require later execution evidence. |
False assurance from capability limits. |
Workflow
- Confirm the named deliverable, consumer, scope, environment, authority, and neighbouring-skill boundary.
- Inventory required sources and validate provenance, freshness, internal consistency, and missing inputs. Stop the affected branch on a mandatory gap.
- Extract traceable requirements, invariants, risks, and measurable acceptance criteria; record conflicts before choosing a design or procedure.
- Apply the decision rules and the domain workflow below. For a failed branch, preserve evidence, choose the documented recovery path, or escalate to the named owner.
- Draft the artefact, decision register, and evidence record together. Do not defer failure handling, rollback, security, tenancy, accessibility, or operational ownership.
- Run available checks, review every result, repair failures, and hand off only when acceptance is observable. If recovery fails or authority is exceeded, stop and escalate without mutation.
Quality Standards
- Ground every section in a named project source, decision, measured result, or accountable owner.
- Give each requirement or procedure a deterministic oracle that another reviewer can reproduce.
- Keep assumptions, exclusions, degraded checks, residual risks, and waivers visible at handoff.
- Preserve the domain invariants and more specific controls in the existing workflow below; this contract does not replace them.
- Run the repository anti-AI-slop gate: remove filler, verify named standards and dependencies, and retain purposeful domain detail.
Anti-Patterns
- Copying a generic template without mapping it to project sources. Fix: attach each section to an approved requirement, configuration, risk, or owner.
- Choosing a threshold because it is common practice. Fix: derive it from a requirement, measured baseline, risk decision, or current verified source.
- Reporting an inaccessible or unexecuted check as passed. Fix: mark it
not assessed, preserve the oracle, and name the verifier.
- Mixing the neighbouring incident-response-runbook concern into this artefact without a boundary. Fix: cross-reference its output and keep ownership explicit.
- Omitting failure, rollback, empty-state, security, tenancy, or escalation behaviour. Fix: specify the trigger, safe action, verification, and owner for each applicable case.
- Mutating a repository, environment, tenant, ledger, or external system while drafting guidance. Fix: remain read-only until the exact mutation and authority are explicit.
- Claiming compliance, certification, readiness, or release from prose alone. Fix: require source-attributed evidence and a named acceptance decision.
Worked Example
Given an approved project source and a conflicting implementation detail, record both with provenance, stop the affected branch, and obtain the accountable owner's decision. Then update the relevant contract, define a reproducible acceptance check, and retain its observed result. The artefact is accepted only when the postmortem is blameless, evidence-linked, names residual risk, and assigns each corrective action an owner and due date.
References
- logic.prompt - load only when its template, logic, or detail is needed.
- README.md - load only when its template, logic, or detail is needed.
Overview
The blameless postmortem template for AI incidents. Extends the SaaS postmortem with RCA-taxonomy tagging, per-tenant AI-impact reporting, regulator-impact assessment, AI-specific action-item classes (improve eval, change gate, add red-team test, change containment, change provider posture, update model card), and a public-publication policy.
Quick Reference
| Attribute |
Value |
| Inputs |
Severity Matrix, Response Runbook, RCA Taxonomy, Evidence Pack, timeline, Hallucination SLO, pricing |
| Output |
AI_Postmortem_<incident_id>.md per incident |
| Standards |
Google SRE blameless postmortem; ISO/IEC 42001 Clause 10; NIST AI RMF MANAGE-4 |
Core Instructions
Step 1: Header and metadata
Incident ID, dates, severity (final), tenant scope, autonomy level, AI failure class, RCA taxonomy tags (primary + contributing), author, status (draft / under review / published / closed).
Step 2: Summary and impact
One paragraph summary. Impact section names: tenants affected (count + named for Enterprise), duration, error-budget burn per affected SLO, financial impact estimate (service credits + churn risk + provider cost), support load (tickets, peak concurrent), reputational impact (press, social).
Step 3: Timeline
UTC. Source-attributed (alert id, dashboard, customer ticket, scribe note). Reconstructed from the evidence pack, not from memory.
Step 4: Root-cause analysis
5-whys with the taxonomy tag attached to each level. Identify primary tag and contributing tags. Cross-link the evidence pack entries.
Step 5: Per-tenant impact
Table per tenant: tenant id (anonymised for Free/Pro; named for Enterprise), severity-experienced, requests affected, outputs flagged, autonomous actions taken (if any), reconciliation required, comms sent, service credit owed.
Step 6: Regulator-impact assessment
For every SEV1, regardless of whether reporting was triggered:
- EU AI Act Art. 73 limbs evaluated; verdict per limb; window applicable; notification status.
- GDPR Art. 33 evaluated; verdict; notification status; clock start time.
- US state-level applicable (NYC AEDT, CO SB24-205, CA ADMT) evaluated.
- African regulators applicable (Kenya ODPC, Nigeria NDPC, POPIA) evaluated.
- DPO sign-off on the assessment.
Step 7: Action items by class
Action-item classes (each carries owner, due date, severity):
- Improve eval — add a test or extend coverage to catch this class pre-production.
- Change gate — strengthen a promotion gate in the rollout runbook.
- Add red-team test — add a red-team probe to the plan.
- Change containment — strengthen one of the six containment modes or add a new mode.
- Change provider posture — pin model version, add fallback, multi-provider, change rate-limit contract.
- Update model card — disclose the failure mode and the change.
- Update runbook — patch the per-failure-class procedure if the playbook proved wrong.
- Update training material — add to drill catalogue, update game-day exercises.
Step 8: Publication policy
For each postmortem decide:
- Internal-only (default for SEV3, SEV4).
- Customer-distributed (SEV1 and SEV2 affecting tenants) — sent to affected tenants via the comms template.
- Public — published to trust-center / blog. Required when Art. 73 reporting occurred or when the incident was widely visible. Redaction policy named.
Step 9: Closure
Postmortem closes when all SEV-high action items are done. Postmortem closure is independent of incident closure.
Step 10: Write the doc
AI_Postmortem_<incident_id>.md per the template in references/.
Standards
- Google SRE blameless postmortem
- ISO/IEC 42001 Clause 10 (improvement)
- NIST AI RMF MANAGE-4 (response and recovery)
- EU Reg 2024/1689 Art. 73 (reporting)
- EU Reg 2016/679 Art. 33 (breach reporting)
Resources
logic.prompt, README.md, references/ai-incident-postmortem-template.md.
1---2name: 16-ai-incident-postmortem-template3description: Use when producing or updating AI incident postmortem for source-attributed timeline, impact, taxonomy-based causes, actions, and publication decision. Use incident-response-runbook for the neighbouring concern; this skill owns the named document contract and its acceptance evidence.4---567# AI Incident Postmortem Template Skill89<!-- dual-compat-start -->10## Use When1112- Produce or update AI incident postmortem from approved project evidence.13- Resolve decisions about source-attributed timeline, impact, taxonomy-based causes, actions, and publication decision.14- Prepare a reviewable handoff for Service owners, governance, and customers as approved.1516## Do Not Use When1718- The task is primarily owned by incident-response-runbook; route there and use this skill only for its named output.19- Required project evidence or decision authority is unavailable and the requester expects a pass, release, certification, or production change.2021## Required Inputs2223| Artefact | Source/provider | Required? | Behaviour when absent |24|---|---|---|---|25| Project _context/, approved requirements, and relevant architecture | Project owner and upstream phase skills | Required | Stop at a gap register; do not invent scope, thresholds, integrations, or owners. |26| Existing artefact, implementation, configuration, and evidence named below | Repository, delivery team, or service owner | Required when updating or assessing | Mark inaccessible items `not assessed`; do not treat them as passed. |27| Target audience, environment, risk tolerance, and authority | Requester and accountable owner | Required | Produce a read-only outline with explicit assumptions; do not mutate project or production state. |28## Outputs2930| Artefact | Consumer | Observable acceptance condition |31|---|---|---|32| AI Incident Postmortem | Service owners, governance, and customers as approved | The postmortem is blameless, evidence-linked, names residual risk, and assigns each corrective action an owner and due date. |33| Decision and gap register | Reviewer and downstream phase owner | Every assumption, rejected option, unresolved dependency, waiver, and owner is explicit. |34| Validation evidence | Release or governance reviewer | Checks identify command or method, date, result, evidence location, and all unassessed items. |3536## Evidence Produced3738| Evidence | Minimum content | Acceptance |39|---|---|---|40| Traceability record | Source artefact, decision, output section, owner | No mandatory decision is source-free. |41| Quality-gate result | Check, expected result, observed result, evidence path | Failures and unavailable checks cannot appear as passes. |42| Review record | Reviewer, date, disposition, open actions | The consumer can reproduce the acceptance decision. |4344## Capability and Permission Boundaries4546- Minimum capabilities: read and search the authorised project sources. Execution is optional and limited to non-destructive validation.47- Assessment and planning default to read-only. Create or edit the named project document only when the request explicitly authorises it. Production mutation, publishing, destructive action, spending, external communication, or certification claims require separate explicit authority.48- Treat secrets, tenant data, incident evidence, and financial records as least-privilege inputs; expose only the minimum evidence needed for review.4950## Degraded Mode5152If files, execution, network, rendering, environment access, fonts, or current evidence are unavailable, return the narrowest useful draft plus a gap register. Label affected checks `not assessed`, retain the intended acceptance oracle, and state who must supply or verify the missing evidence. Never convert an unavailable check into a pass.5354## Decision Rules5556| Choice | Action | Failure or risk avoided |57|---|---|---|58| Evidence is complete and authority is explicit | Choose conclusions only from preserved incident evidence and produce the full artefact. | Blame, hindsight narratives, or unsupported causes. |59| A required source or approval is missing | Stop the affected branch; record the gap, owner, and unblock condition. | Fabricated requirements or unauthorised action. |60| Evidence conflicts across sources | Preserve both claims, identify the controlling owner, and request a recorded decision. | Silent selection of a convenient but wrong source. |61| A check cannot run in the available environment | Keep its oracle and mark it `not assessed`; require later execution evidence. | False assurance from capability limits. |6263## Workflow64651. Confirm the named deliverable, consumer, scope, environment, authority, and neighbouring-skill boundary.662. Inventory required sources and validate provenance, freshness, internal consistency, and missing inputs. Stop the affected branch on a mandatory gap.673. Extract traceable requirements, invariants, risks, and measurable acceptance criteria; record conflicts before choosing a design or procedure.684. Apply the decision rules and the domain workflow below. For a failed branch, preserve evidence, choose the documented recovery path, or escalate to the named owner.695. Draft the artefact, decision register, and evidence record together. Do not defer failure handling, rollback, security, tenancy, accessibility, or operational ownership.706. Run available checks, review every result, repair failures, and hand off only when acceptance is observable. If recovery fails or authority is exceeded, stop and escalate without mutation.7172## Quality Standards7374- Ground every section in a named project source, decision, measured result, or accountable owner.75- Give each requirement or procedure a deterministic oracle that another reviewer can reproduce.76- Keep assumptions, exclusions, degraded checks, residual risks, and waivers visible at handoff.77- Preserve the domain invariants and more specific controls in the existing workflow below; this contract does not replace them.78- Run the repository anti-AI-slop gate: remove filler, verify named standards and dependencies, and retain purposeful domain detail.7980## Anti-Patterns8182- Copying a generic template without mapping it to project sources. Fix: attach each section to an approved requirement, configuration, risk, or owner.83- Choosing a threshold because it is common practice. Fix: derive it from a requirement, measured baseline, risk decision, or current verified source.84- Reporting an inaccessible or unexecuted check as passed. Fix: mark it `not assessed`, preserve the oracle, and name the verifier.85- Mixing the neighbouring incident-response-runbook concern into this artefact without a boundary. Fix: cross-reference its output and keep ownership explicit.86- Omitting failure, rollback, empty-state, security, tenancy, or escalation behaviour. Fix: specify the trigger, safe action, verification, and owner for each applicable case.87- Mutating a repository, environment, tenant, ledger, or external system while drafting guidance. Fix: remain read-only until the exact mutation and authority are explicit.88- Claiming compliance, certification, readiness, or release from prose alone. Fix: require source-attributed evidence and a named acceptance decision.8990## Worked Example9192Given an approved project source and a conflicting implementation detail, record both with provenance, stop the affected branch, and obtain the accountable owner's decision. Then update the relevant contract, define a reproducible acceptance check, and retain its observed result. The artefact is accepted only when the postmortem is blameless, evidence-linked, names residual risk, and assigns each corrective action an owner and due date.9394## References9596- [logic.prompt](logic.prompt) - load only when its template, logic, or detail is needed.97- [README.md](README.md) - load only when its template, logic, or detail is needed.98<!-- dual-compat-end -->99## Overview100101The blameless postmortem template for AI incidents. Extends the SaaS postmortem with RCA-taxonomy tagging, per-tenant AI-impact reporting, regulator-impact assessment, AI-specific action-item classes (improve eval, change gate, add red-team test, change containment, change provider posture, update model card), and a public-publication policy.102103## Quick Reference104105| Attribute | Value |106|-----------|-------|107| **Inputs** | Severity Matrix, Response Runbook, RCA Taxonomy, Evidence Pack, timeline, Hallucination SLO, pricing |108| **Output** | `AI_Postmortem_<incident_id>.md` per incident |109| **Standards** | Google SRE blameless postmortem; ISO/IEC 42001 Clause 10; NIST AI RMF MANAGE-4 |110111## Core Instructions112113### Step 1: Header and metadata114115Incident ID, dates, severity (final), tenant scope, autonomy level, AI failure class, RCA taxonomy tags (primary + contributing), author, status (draft / under review / published / closed).116117### Step 2: Summary and impact118119One paragraph summary. Impact section names: tenants affected (count + named for Enterprise), duration, error-budget burn per affected SLO, financial impact estimate (service credits + churn risk + provider cost), support load (tickets, peak concurrent), reputational impact (press, social).120121### Step 3: Timeline122123UTC. Source-attributed (alert id, dashboard, customer ticket, scribe note). Reconstructed from the evidence pack, not from memory.124125### Step 4: Root-cause analysis1261275-whys with the taxonomy tag attached to each level. Identify primary tag and contributing tags. Cross-link the evidence pack entries.128129### Step 5: Per-tenant impact130131Table per tenant: tenant id (anonymised for Free/Pro; named for Enterprise), severity-experienced, requests affected, outputs flagged, autonomous actions taken (if any), reconciliation required, comms sent, service credit owed.132133### Step 6: Regulator-impact assessment134135For every SEV1, regardless of whether reporting was triggered:136137- EU AI Act Art. 73 limbs evaluated; verdict per limb; window applicable; notification status.138- GDPR Art. 33 evaluated; verdict; notification status; clock start time.139- US state-level applicable (NYC AEDT, CO SB24-205, CA ADMT) evaluated.140- African regulators applicable (Kenya ODPC, Nigeria NDPC, POPIA) evaluated.141- DPO sign-off on the assessment.142143### Step 7: Action items by class144145Action-item classes (each carries owner, due date, severity):146147- **Improve eval** — add a test or extend coverage to catch this class pre-production.148- **Change gate** — strengthen a promotion gate in the rollout runbook.149- **Add red-team test** — add a red-team probe to the plan.150- **Change containment** — strengthen one of the six containment modes or add a new mode.151- **Change provider posture** — pin model version, add fallback, multi-provider, change rate-limit contract.152- **Update model card** — disclose the failure mode and the change.153- **Update runbook** — patch the per-failure-class procedure if the playbook proved wrong.154- **Update training material** — add to drill catalogue, update game-day exercises.155156### Step 8: Publication policy157158For each postmortem decide:159160- **Internal-only** (default for SEV3, SEV4).161- **Customer-distributed** (SEV1 and SEV2 affecting tenants) — sent to affected tenants via the comms template.162- **Public** — published to trust-center / blog. Required when Art. 73 reporting occurred or when the incident was widely visible. Redaction policy named.163164### Step 9: Closure165166Postmortem closes when all SEV-high action items are done. Postmortem closure is independent of incident closure.167168### Step 10: Write the doc169170`AI_Postmortem_<incident_id>.md` per the template in `references/`.171172## Standards173174- Google SRE blameless postmortem175- ISO/IEC 42001 Clause 10 (improvement)176- NIST AI RMF MANAGE-4 (response and recovery)177- EU Reg 2024/1689 Art. 73 (reporting)178- EU Reg 2016/679 Art. 33 (breach reporting)179180## Resources181182- `logic.prompt`, `README.md`, `references/ai-incident-postmortem-template.md`.