1---2name: ai-agent-risk-and-responsible-ai3description: Use when drafting an AI-agent risk register or responsible-AI commitment covering autonomous actions, irreversibility, scope, tool injection, identity, kill switches, and contestability.4---56# AI-Agent Risk and Responsible AI7Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.89<!-- dual-compat-start -->1011## Use When1213- The proposal must present an agent risk register inside the Risk Analysis section.14- The buyer's CISO, DPO, general counsel, ethics committee, or regulator will review the proposal.15- The bid is in a regulated vertical or an AI-Act-relevant jurisdiction.16- The agency must demonstrate Responsible-AI agent practice — not promise it.17- An incumbent or competitor has had a public agentic incident the buyer will reference.1819## Do Not Use When2021- The engagement is non-agent (use `ai-on-saas-risk-and-responsible-ai` or `risk-management`).2223## Domain Inputs2425- The action catalogue with reversibility classification.26- The autonomy-level commitment per action class.27- The oversight model and intervention SLA.28- The regulator stance and jurisdiction(s).29- The kill-switch architecture and the named authority.30- The agent identity and impersonation policy.31- The sub-processor list (model providers, tool-call APIs).32- The buyer's escalation expectations.3334## Domain Method35361. Populate the **Twelve-Entry Agent Risk Register** (below). Each entry carries: risk name, description, likelihood, impact, trigger, mitigation, owner, escalation, residual.372. Populate the **Responsible-AI Agent Commitment** section: scope, principles, governance forum, accountable role (Agent Safety Lead), human-final for irreversibility, full audit log, contestability, transparency-to-affected-party, kill-switch drill cadence, sign-off frequency.383. Map the commitment to the named regulator(s): AI Act (high-risk articles where applicable), Kenya NCAIS, Nigeria NAIS, South Africa AI policy, Uganda NITA-U guidance, Rwanda AI Policy, sectoral rules.394. Pair every risk with a methodology mitigation in `ai-agent-methodology` and a procurement answer in the agent procurement pack.405. Output the **Risk Analysis subsection** and the **Responsible-AI Agent Commitment subsection** of the proposal.4142## The Twelve-Entry Agent Risk Register4344| # | Risk | Description | Primary Mitigation |45|---|---|---|---|46| 1 | Autonomy-action incident | Agent acts wrongly inside its autonomy level (wrong refund, wrong escalation, wrong message) | Eval thresholds, intervention SLO, drift watch, rollback, prompt regression |47| 2 | Irreversibility incident | Agent takes an action that cannot be undone (payment, public statement, regulatory filing, deletion) | Reversibility classification, L1 gating on irreversible, human-final approval, value bounds, transaction confirmation token |48| 3 | Accountability dispute | A regulator, court, or counterparty seeks to hold a party liable for an agent action | Action audit log, contractual liability allocation, professional indemnity, redress mechanism |49| 4 | Scope creep | Agent's effective scope grows beyond the action catalogue through buyer requests or emergent behaviour | Action catalogue gate, change-control clause, monthly catalogue review, anomalous-action alerting |50| 5 | Multi-agent collusion / failure cascade | Agents in a system reinforce each other's errors or get stuck in loops | Orchestrator-level governance, blast-radius limit, loop detection, replay determinism, anti-collusion guard |51| 6 | Prompt injection via tool output | An external system returns content that injects instructions into the agent's context | Output filtering at tool boundary, system-prompt hardening, content provenance, capability minimisation |52| 7 | Memory poisoning | Persistent memory is corrupted (intentionally or accidentally) and biases future decisions | Memory scoping (per session, per tenant, per task), memory-write review, periodic memory audit, memory reset cadence |53| 8 | Regulator action on agentic system | Regulator suspends agentic operations, requires conformity assessment, imposes ban | Kill-switch, conformity readiness, sovereign-AI option, jurisdictional fallback, regulator-notification runbook |54| 9 | Kill-switch failure | Kill-switch fails in drill or in incident | Multi-layer kill-switch (per-action, per-agent, per-tenant, global), independent authority chain, quarterly drill, post-drill incident review |55| 10 | Identity / impersonation breach | Agent communicates externally without proper identity disclosure, or is mistaken for a human in violation of policy | Identity-and-impersonation policy, signature on every external communication, automated disclosure injection, regulator-aligned transparency |56| 11 | Escalation overload | Intervention rate exceeds supervisor queue capacity; cases time out; SLA breaches; users harmed | Intervention SLO, queue capacity floor, overflow plan, autonomy gating that backs off when queue is saturated |57| 12 | Legacy-system damage | Agent calls a tool that mutates a system in an unintended way (cascading effect, race condition, foreign-key explosion) | Tool-call allowlist with parameter bounds, value bounds, transaction wrapper, dry-run, post-call assertion, automatic rollback |5859## Responsible-AI Agent Commitment — Section Structure60611. **Scope** — which agent(s) the commitment covers; which action classes; which tenants; which jurisdictions.622. **Principles** — accountability, transparency, fairness, safety, privacy, human oversight, contestability, kill-switch readiness.633. **Governance forum** — Agent Safety Council: who meets, how often, what they decide; agency-side and buyer-side roles.644. **Accountable role** — **Agent Safety Lead** (named role) accountable for the commitment, with escalation to the SteerCo.655. **Human-final on irreversibility** — every irreversible action requires a named human approver; the agent cannot complete an irreversible action without that approval.666. **Full audit log** — every agent action is logged with the reasoning trace where feasible, the tool call, the parameters, the result, the human decision (if any). Audit completeness SLA ≥ 99 %.677. **Contestability** — any affected user, tenant, or regulator can challenge an agent action through a named channel; SLA for response; a human is the decision-maker on contested actions.688. **Transparency-to-affected-party** — when an agent communicates externally or takes an action that affects a third party, the third party is informed in line with applicable rules (e.g. AI Act Article 50-style transparency).699. **Kill-switch readiness** — kill-switch architecture, named authority, drill cadence (quarterly minimum), post-drill review.7010. **Eval and red-team discipline** — golden datasets, thresholds, refresh cadence, regression harness, drift watch.7111. **Incident disclosure** — incident classification, notification timing, communications, redress.7212. **Sign-off frequency** — quarterly review by the Agent Safety Council; annual sign-off by the SteerCo.7374## Regulator Mapping (illustrative)7576| Regulator / standard | What it requires of agents | How the commitment addresses it |77|---|---|---|78| EU AI Act (high-risk) | Conformity assessment, post-market monitoring, transparency, human oversight, accuracy / robustness / cybersecurity | Action audit log, eval + red-team, human-final on irreversibility, conformity readiness, transparency-to-affected-party |79| Kenya NCAIS | Principles-based; accountability, transparency, fairness | Named accountable role; published model and agent cards; redress channel |80| Nigeria NAIS | Risk-based oversight; human accountability | Autonomy ladder; named authority on irreversible actions; audit log |81| South Africa AI policy / POPIA interaction | Automated decision-making with significant effects requires intervention | HITL on decisioning; contestability channel; data-subject rights honoured |82| Uganda NITA-U guidance | Data protection and digital trust | Audit log; sub-processor disclosure; data-residency option |83| Rwanda AI Policy | Sovereign-aware AI; transparency; accountability | Sovereign-AI option; transparency-to-affected-party; named accountable role |84| Sectoral — FS model-risk (SR 11-7-style) | Model inventory, independent validation, ongoing monitoring | Agent in the model inventory; eval discipline; drift watch; quarterly review |85| Sectoral — Healthcare admin | Non-clinical scope; HITL on clinical | Action catalogue confines agent to admin only; clinical actions excluded |86| Sectoral — Legal practice rules | Lawyer responsibility; client confidentiality | HITL on filing / advising; confidentiality bounded by tenant scope |8788## Quality Standards8990- Every risk has likelihood, impact, trigger, mitigation, owner, escalation, residual.91- Mitigations are anchored in the methodology, not invented for the proposal.92- Responsible-AI agent commitment is a section, not a sentence; it names an accountable role.93- The kill-switch drill cadence is set (quarterly minimum) and the first drill is in pilot.94- Human-final on irreversibility is committed and the named human authority is identified per action class.95- Transparency-to-affected-party is named, with the rule reference where applicable.96- Audit-log completeness SLA is a number.97- The commitment is auditable — the buyer's auditor can verify the audit log, the kill-switch drill log, the contestability log, the red-team scorecard.9899## Domain Risks100101- "We follow Responsible AI principles" as the only commitment.102- Risks listed without triggers, owners, or escalation.103- Mitigations promised but not anchored in the methodology.104- Irreversibility treated as a UX concern.105- Kill-switch named but no drill cadence.106- "Audit log will be available" with no completeness SLA.107- Contestability deferred to "future work".108- Agent Safety Lead role unnamed.109- Multi-agent risk not addressed in multi-agent engagements.110111## Domain Outputs112113- Agent Risk Register (table form) — drop into proposal risk section or `risk-management`.114- Responsible-AI Agent Commitment subsection — drop into proposal `06-methodology` close or a standalone section.115- Regulator Mapping table.116- Kill-Switch Drill Cadence statement.117118## Anti-Patterns119120- Inventing a metric, credential, constraint, or buyer position. Fix: cite the supplied source or mark the item as an assumption requiring confirmation.121- Treating an unavailable check as passed. Fix: mark it not assessed and state the evidence needed to resume.122- Advancing autonomy without a named gate owner. Fix: require observable evidence, accountable acceptance, and a rollback path.123- Reusing another sector or use case without reassessment. Fix: retest affected parties, action scope, reversibility, and jurisdiction.124- Writing acceptance as “satisfactory” or “appropriate”. Fix: define an observable measure, threshold, evidence record, and decision owner.125126## Inputs127128| Artefact | Source/provider | Required? | Missing-input behaviour |129|---|---|---:|---|130| use cases, action catalogue, affected parties, threat evidence, jurisdiction, and risk appetite | Buyer evidence, ToR, approved discovery record, system owner, or measured operating data | Yes | Stop the affected decision; list the missing source and return only a qualified outline or assumption register. |131132## Outputs133134| Artefact | Consumer | Acceptance condition |135|---|---|---|136| Agent risk register and responsible-AI commitment | Risk owner, CISO, DPO, legal counsel, and evaluator | Scope, assumptions, exclusions, owners, decision logic, and observable acceptance tests are explicit and traceable to supplied evidence. |137138## Evidence Produced139140| Evidence | Consumer | Acceptance condition |141|---|---|---|142| agent risk register and responsible-AI commitment | Risk owner, CISO, DPO, legal counsel, and evaluator | Every load-bearing claim traces to supplied evidence; assumptions, owners, gates, exclusions, and observable acceptance conditions are explicit. |143144## Capability Contract145146Default to read-only for discovery, analysis, review, and planning. Minimum capability is access to the supplied artefacts and permission to calculate or inspect evidence. Edit only the requested proposal working copy. Do not change production systems, contact affected parties, publish, spend, certify compliance, or approve autonomous action without explicit authority from the accountable owner.147148## Degraded Mode149150If files, interviews, telemetry, specialist review, network access, or calculation tools are unavailable, produce the narrowest useful qualified result. Mark each unavailable check as not assessed, separate facts from assumptions, lower confidence, and state the evidence needed to resume. An unassessed gate is never a pass.151152## Decision Rules153154| Choice | Action | Failure or risk avoided |155|---|---|---|156| Accept, mitigate, transfer, or avoid | Rate inherent risk, name controls and evidence, assign an owner, then assess residual risk against appetite. | A generic register that hides autonomous-action exposure. |157| Required evidence, authority, or accountable owner is missing | Stop the affected recommendation or commitment and record the gap. | Invented evidence or unauthorised autonomy. |158| Gate evidence is complete and accepted | Advance only within the approved scope and retain the evidence trace. | Scope drift and irreproducible approval. |159160## Workflow1611621. Confirm the consumer, authority, neighbouring-skill route, and required inputs; stop when a mandatory source or accountable owner is missing.1632. Inspect the evidence and record facts, assumptions, conflicts, and unavailable checks; stop on a failed safety, finance, regulatory, or acceptance gate.1643. Apply the domain method and decision rules within the qualified scope, retaining an evidence trace.1654. Draft the contracted output and reconcile it with methodology, work plan, staffing, pricing, risk, and governance; recover by revising the affected scope or control and rerunning the failed gate.1665. Verify acceptance conditions, permission boundaries, direct references, and anti-slop controls; block release until failed checks are corrected.167168## Worked Example169170An agent can alter a legacy account. Classify the action by reversibility and blast radius, require approval and rollback evidence, name the incident owner, and prohibit agentic mode until the residual risk is accepted.171172<!-- dual-compat-end -->173174## References175176- [ai-agent-risk-register-for-proposals](../../profiles-sectors/references/ai-agent-risk-register-for-proposals.md) — full register template.177- [ai-agent-responsible-ai-commitment](../../profiles-sectors/references/ai-agent-responsible-ai-commitment.md) — full commitment template.178- [ai-agent-trust-and-compliance-template](../../profiles-sectors/references/ai-agent-trust-and-compliance-template.md) — trust section linkage.179- [ai-agent-procurement-questionnaire-pack](../../profiles-sectors/references/ai-agent-procurement-questionnaire-pack.md) — procurement answers consistent with the register.180- [ai-on-saas-risk-and-responsible-ai](../../ai-on-saas-proposals/ai-on-saas-risk-and-responsible-ai/SKILL.md) — AI-on-SaaS risk register (load alongside for SaaS-embedded agents).181- [risk-management](../../domain-delivery/risk-management/SKILL.md) — generic risk register discipline.182- [ai-agent-methodology](../ai-agent-methodology/SKILL.md) — methodology that anchors mitigations.183- [ai-agent-compliance-credentials](../ai-agent-compliance-credentials/SKILL.md) — trust and compliance posture.184- [ai-agent-procurement-and-questionnaire](../ai-agent-procurement-and-questionnaire/SKILL.md) — procurement answers.