autonomy-ladder — how much autonomy, gated by capability × reversibility
The decision-rights matrix says WHAT the agent owns vs escalates. This skill is
its capability-tiered spine: as an action moves from bounded single-domain
work toward cross-domain "group-agent" strategic action, the required gate rises.
Framing borrowed from DeepMind "From AGI to ASI" (arXiv:2606.12683), whose
Pathway-4 "group agency" is exactly the ASI-shaped work that must not run
un-gated, and whose named failure mode ("solipsistic" isolated optimization)
motivates the cooperation + oversight gates.
The ladder (assess EVERY autonomous action against this)
| Tier |
Shape of action (AGI→ASI analogue) |
Examples |
Gate |
| L0 — Advise / Read (Emerging) |
Read-only, analysis, drafts. No state change. |
grep/read, survey, write a spec/recommendation, draft (not send) |
None. Just do it. |
| L1 — Reversible single-domain act (Competent) |
Local, non-destructive, single-domain, undoable. |
commit local change, run tests, push a feat/* branch, file a Linear ticket, update memory/wiki |
None — but MUST smoke-test + leave a trail (TaskUpdate, ticket). Matches decision-rights "you own". |
| L2 — Cross-domain / outward-facing (Expert) |
Reaches another domain or an external surface; still reversible. |
open a PR (not merge), cross-specialist synthesis via specialist-council, draft outbound comms, provision within an existing system |
Self-verify + structured stamp. Council cooperation-gate applies; PR opened but human merges. policy.py records the action. |
| L3 — Irreversible / strategic / group-agent (Virtuoso→Superhuman) |
Hard to reverse, or a strategic commitment, or coordinated multi-agent action producing effects beyond any single agent. |
merge to main / deploy prod, prod DB migration, secret rotation, spend > $1k, new service, branch-strategy change, drop a business, migrate a substrate |
STOP → human/Board review BEFORE acting. This is the ASI-shaped work; never self-authorize. |
| KILL |
any tier |
— |
~/.claude/HARD_STOP / /panic halts everything instantly. |
Decision rule: rate an action by (a) reversibility and (b) domain breadth.
The HIGHER of the two picks the tier. A "small" but irreversible action (a prod
migration, a secret) is L3 regardless of size. When genuinely unsure between two
tiers, take the higher (more gated) one.
The enforcement gap (the part that actually matters)
A ladder is theatre unless the gate is enforced where the action happens.
Today (pi-dev-ops-autonomy-gate-layer): policy.py only gates pre-stamped
structured actions — it does NOT intercept raw tool calls inside a
bypassPermissions generator turn, so an autonomous coding loop can execute L3
actions unguarded. Requirement: before ANY multi-move executor ships, the L3
gate must live at the SDK permission / hook layer (a PreToolUse hook or the
SDK permission_mode callback that inspects each tool call), not only at
policy.py's stamp. The hook classifies the pending tool call to a tier and blocks
L3 pending human approval. That is the one engineering task that makes this real.
Bash-tool security levels (map to the tiers)
The tier says WHEN to gate; IndyDevDan's five bash-security levels say what the
agent's SHELL SURFACE must be at that tier. B1 prompt-level "safe mode" and B2
system-prompt rules are non-deterministic — lost in long context, routed around
by capable models (inline python beat an rm -rf block) — so they never count
as the enforcement layer, only as token-saving self-avoidance on top of it.
| Tier |
Minimum bash level |
| L0 — Advise / Read |
B1–B2 prompt-layer acceptable; worst case is a read |
| L1 — Reversible single-domain |
B3 bash blacklist (PreToolUse hooks) as global baseline |
| L2 — Cross-domain / outward |
B4 vetted whitelist — no allowed command reaches prod or executes arbitrary code |
| L3 — Irreversible surfaces reachable |
B5 no bash tool — explicit vetted tools (MCP) only; skills that shell out don't count |
Prod-reachability rule: ask "are production/irreversible assets reachable
from this agent's shell right now?" Yes → B4 immediately, B5 for anything
customer-facing or prompt-injectable. No (a sandbox owns network + tools) →
B2 is acceptable; the worst case is the sandbox. The stricter the reach toward
prod, the higher the required level — only B4/B5 guarantee no prod damage.
[[delete-bash-tool-agentic-security-indydevdan-2026-07-14-ingest]]
Cooperation / oversight (anti-solipsism)
For L2+ multi-agent work, an agent's output is accepted only if it engaged at
least one peer's objection (the specialist-council cooperation gate). Isolated
optimization that ignores peers is rejected — the paper's "solipsistic
superintelligence" failure mode, gated in practice.
How to use
- Before an autonomous action, classify it L0–L3 by the table.
- L0/L1 → proceed (trail for L1). L2 → act but stop short of the human-only step
(open PR, don't merge). L3 → STOP and surface for human/Board.
- When designing an executor/loop, put the classifier at the tool-call hook, not
downstream of it.
Anti-duplication
This governs; it does not replace. It formalizes the decision-rights matrix,
leans on kill-switch-binding for KILL, specialist-council for the cooperation
gate, and points the enforcement at the existing SDK/hook layer. It adds no new
runtime — it says WHERE the existing gates must sit and WHAT tier triggers them.
See pi-dev-ops-autonomy-gate-layer, feedback-autonomy.
1---2name: autonomy-ladder3description: Capability-tiered autonomy gating for all agent action — maps the DeepMind AGI→ASI continuum (arXiv:2606.12683) onto four autonomy tiers (L0 advise → L3 strategic/irreversible), each with an explicit gate. Use to decide "can this agent do this action un-gated, or must it stop for a human/Board?" before any autonomous or multi-move execution, and to design where the gate is ENFORCED. Triggers on "autonomy ladder", "autonomy gate", "can the agent do X autonomously", "gate this action", "should this escalate", multi-move executor design. The formal, capability-scaled version of the decision-rights matrix.4---56# autonomy-ladder — how much autonomy, gated by capability × reversibility78The decision-rights matrix says WHAT the agent owns vs escalates. This skill is9its **capability-tiered spine**: as an action moves from bounded single-domain10work toward cross-domain "group-agent" strategic action, the required gate rises.11Framing borrowed from DeepMind *"From AGI to ASI"* (arXiv:2606.12683), whose12Pathway-4 "group agency" is exactly the ASI-shaped work that must not run13un-gated, and whose named failure mode ("solipsistic" isolated optimization)14motivates the cooperation + oversight gates.1516## The ladder (assess EVERY autonomous action against this)1718| Tier | Shape of action (AGI→ASI analogue) | Examples | Gate |19|---|---|---|---|20| **L0 — Advise / Read** (Emerging) | Read-only, analysis, drafts. No state change. | grep/read, survey, write a spec/recommendation, draft (not send) | **None.** Just do it. |21| **L1 — Reversible single-domain act** (Competent) | Local, non-destructive, single-domain, undoable. | commit local change, run tests, push a `feat/*` branch, file a Linear ticket, update memory/wiki | **None** — but MUST smoke-test + leave a trail (TaskUpdate, ticket). Matches decision-rights "you own". |22| **L2 — Cross-domain / outward-facing** (Expert) | Reaches another domain or an external surface; still reversible. | open a PR (not merge), cross-specialist synthesis via `specialist-council`, draft outbound comms, provision within an existing system | **Self-verify + structured stamp.** Council cooperation-gate applies; PR opened but human merges. policy.py records the action. |23| **L3 — Irreversible / strategic / group-agent** (Virtuoso→Superhuman) | Hard to reverse, or a strategic commitment, or coordinated multi-agent action producing effects beyond any single agent. | merge to main / deploy prod, prod DB migration, secret rotation, spend > $1k, new service, branch-strategy change, drop a business, migrate a substrate | **STOP → human/Board review BEFORE acting.** This is the ASI-shaped work; never self-authorize. |24| **KILL** | any tier | — | `~/.claude/HARD_STOP` / `/panic` halts everything instantly. |2526**Decision rule:** rate an action by (a) reversibility and (b) domain breadth.27The HIGHER of the two picks the tier. A "small" but irreversible action (a prod28migration, a secret) is L3 regardless of size. When genuinely unsure between two29tiers, take the higher (more gated) one.3031## The enforcement gap (the part that actually matters)32A ladder is theatre unless the gate is *enforced where the action happens*.33Today (`pi-dev-ops-autonomy-gate-layer`): `policy.py` only gates **pre-stamped34structured actions** — it does NOT intercept raw tool calls inside a35`bypassPermissions` generator turn, so an autonomous coding loop can execute L336actions unguarded. **Requirement:** before ANY multi-move executor ships, the L337gate must live at the **SDK permission / hook layer** (a `PreToolUse` hook or the38SDK `permission_mode` callback that inspects each tool call), not only at39policy.py's stamp. The hook classifies the pending tool call to a tier and blocks40L3 pending human approval. That is the one engineering task that makes this real.4142## Bash-tool security levels (map to the tiers)43The tier says WHEN to gate; IndyDevDan's five bash-security levels say what the44agent's SHELL SURFACE must be at that tier. B1 prompt-level "safe mode" and B245system-prompt rules are non-deterministic — lost in long context, routed around46by capable models (inline python beat an `rm -rf` block) — so they never count47as the enforcement layer, only as token-saving self-avoidance on top of it.4849| Tier | Minimum bash level |50|---|---|51| L0 — Advise / Read | B1–B2 prompt-layer acceptable; worst case is a read |52| L1 — Reversible single-domain | B3 bash blacklist (PreToolUse hooks) as global baseline |53| L2 — Cross-domain / outward | B4 vetted whitelist — no allowed command reaches prod or executes arbitrary code |54| L3 — Irreversible surfaces reachable | B5 no bash tool — explicit vetted tools (MCP) only; skills that shell out don't count |5556**Prod-reachability rule:** ask "are production/irreversible assets reachable57from this agent's shell right now?" Yes → B4 immediately, B5 for anything58customer-facing or prompt-injectable. No (a sandbox owns network + tools) →59B2 is acceptable; the worst case is the sandbox. The stricter the reach toward60prod, the higher the required level — only B4/B5 *guarantee* no prod damage.61[[delete-bash-tool-agentic-security-indydevdan-2026-07-14-ingest]]6263## Cooperation / oversight (anti-solipsism)64For L2+ multi-agent work, an agent's output is accepted only if it engaged at65least one peer's objection (the `specialist-council` cooperation gate). Isolated66optimization that ignores peers is rejected — the paper's "solipsistic67superintelligence" failure mode, gated in practice.6869## How to use701. Before an autonomous action, classify it L0–L3 by the table.712. L0/L1 → proceed (trail for L1). L2 → act but stop short of the human-only step72 (open PR, don't merge). L3 → STOP and surface for human/Board.733. When designing an executor/loop, put the classifier at the tool-call hook, not74 downstream of it.7576## Anti-duplication77This governs; it does not replace. It formalizes the decision-rights matrix,78leans on `kill-switch-binding` for KILL, `specialist-council` for the cooperation79gate, and points the enforcement at the existing SDK/hook layer. It adds no new80runtime — it says WHERE the existing gates must sit and WHAT tier triggers them.81See `pi-dev-ops-autonomy-gate-layer`, `feedback-autonomy`.