# Claude Code Hooks

> How to write, test, register, and debug Claude Code hooks — PreToolUse / PostToolUse / SessionStart / Stop Bash guards that enforce a rule the model would otherwise talk itself past. Use whenever the user wants to create a hook, block/intercept a tool call, turn a repeatedly-violated rule into a hard gate, add a guard rail, debug a hook that misfires or "poisons the session", register a hook across profiles, or mentions hooks / PreToolUse / Stop hook / 拦截 / 守卫 / 钩子 / 拦下. Bakes in the hard-won pitfalls: UserPromptSubmit only ever sees user input, never Claude's own text — a rule about Claude's own output belongs on Stop instead; token-level shlex matching (never awk splitting); bash -n + real-JSON end-to-end testing BEFORE registering (a corrupted PreToolUse hook poisons every Bash call); SSOT + symlink so a ~/.claude reinstall can't lose it; multi-profile convergence; and human-confirmation release gates. Reach for this even for "make it stop doing X" — a durable stop is a hook, not a reminder.

- Skill: `gabrielmoreira/claude-code-hooks` (Agent Skill)
- Install (CLI): `npx skillmds@latest add gabrielmoreira/claude-code-hooks`
- Raw SKILL.md: https://api.skillmd.com/api/skills/gabrielmoreira/claude-code-hooks/raw
- Safety review: pending (external: skill-scanner PASS, skillspector FAIL)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: gabrielmoreira (https://skillmd.com/u/gabrielmoreira)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/gabrielmoreira/claude-code-hooks

---


# Claude Code Hooks

Claude Code fires **hooks** at tool-call boundaries. A hook is a shell command
that receives a JSON event on stdin and, for blocking hooks, decides via its
**exit code** whether the tool call proceeds. This is the only mechanism that
*structurally* stops a behavior — a prose rule in CLAUDE.md is a suggestion the
completion drive can override; a hook is a wall.

## When a hook is the right tool (and when it isn't)

Write a hook when **a rule keeps getting violated even though it's already
written down**. The tell: you added the prose rule, it read clearly, and the
behavior recurred anyway — because at the moment of action, attention is 100%
on "get the thing done" and the reminder loses. That recurrence is the signal
to move the rule from prose (advisory) to a hook (enforced). Governance rule of
thumb: *Tier-0 irreversible action + only prose, no hook → it should be a hook.*
(**Tier-0** here = an action whose damage cannot be undone from inside the session:
destroying uncommitted work, pushing secrets to a remote, deleting files, publishing
something outward. The test is reversibility, not severity.)

Do **not** reach for a hook when: the rule has never actually recurred (don't
pre-build guards for hypothetical mistakes — cost with no proven benefit), or
the "rule" is a judgment call with no mechanical signature (a hook can only
match tokens/patterns; it can't judge whether a design is good).

If the symptom is “it keeps reviewing / waiting / retrying,” do not assume the
answer is another hook. First complete the **Loop Contract** in rule 7 and read
[pitfall #36](references/hook_pitfalls.md#36-a-self-applied-review-rule-can-loop-without-any-hook).
The loop may be created entirely by an agent repeatedly applying a prose rule.

## Hook types and what the exit code means

| Type | Fires | Exit 0 | Exit 2 | Other |
|---|---|---|---|---|
| **PreToolUse** | before a tool runs | allow | **block** the call (stderr → shown to model as guidance) | any other exit = "non-blocking error" → **the call proceeds** — but only while stdout carries no valid JSON. Claude Code reads JSON output on **every** exit code, and valid JSON overrides the code entirely. The skeletons here print nothing on stdout, so their fail-open reasoning holds; add a `permissionDecision` payload and the exit code stops being the decision |
| **PostToolUse** | after a tool ran | quiet **unless it prints a `hookSpecificOutput` JSON on stdout — that is how context injection works, and it happens at exit 0** | feedback to the model (can't un-run the tool) | — |
| **SessionStart** | session begins | proceed | **cannot block** — stderr shows the user a hook-error notice, Claude never sees it, the session starts anyway | **exit 0 anyway**: not because a non-zero would block (it can't), but because anything non-zero puts a `<hook> hook error` in the user's transcript on every single session start. Takes a `matcher` on *how the session started* — `startup`, `resume`, `clear`, `compact`, `fork` |
| **Stop** (+ `SubagentStop`) | the model is about to finish responding | let it stop | **block the stop** — forces the model to keep going (stderr → fed back as the reason) | loop safety: the hook checks `stop_hook_active` (necessary, **not** sufficient — rule 7). The harness's consecutive-block ceiling (default 8) is **not** a general backstop — its counter resets on any continuation that executed tools, so it never arrives for a hook whose remediation involves tool calls, which is most of them (#27). Carry your own bound. All Stop hooks for an event run **in parallel** — one block round can carry several hooks' feedback |

- **PreToolUse** is the workhorse for stopping a *tool call* — the four types in this
  table are the ones this file teaches, not the complete set of blockable events, and
  the **official** hooks reference (docs.claude.com / code.claude.com, not the
  `references/` files in this bundle — those cover only the four types above) now
  lists many more blockable events, including `UserPromptSubmit`, `PreCompact`,
  `TeammateIdle`, and task and config events. If what you need to gate is not a tool
  call, look there before forcing it onto PreToolUse.
  `matcher` selects the tool (`Bash`, `Agent`, `WebFetch`, …) — **and how it is
  matched depends on the characters you use**: a matcher containing only letters,
  digits, `_`, `-`, spaces, `,` and `|` is compared as an **exact string** (or a
  `|`/`,`-separated list of exact strings); anything else is treated as an
  **unanchored JavaScript regex**. Both directions bite silently — `Edit.*` also
  matches `NotebookEdit` (anchor it `^Edit$`), while `mcp__memory` matches **nothing**
  because it is all exact-match characters and no tool is named exactly that (you
  want `mcp__memory__.*`). Matching is case-sensitive. Exit 2 blocks and
  the hook's **stderr** becomes the message the model sees — so put the *why* and
  the *correct alternative* there, not just "blocked".
- **PostToolUse** can't undo, but it can **inject authoritative context** so a
  later hallucination can't stand (e.g. re-read the real git HEAD after a commit
  and surface it — the model can't "believe it committed" against injected truth).
- **SessionStart** is for **health checks of the guard rails themselves** —
  silent when healthy, warn on breakage, always exit 0. Note *why*: it is not that
  a non-zero exit would block the session (it cannot), but that it would print a
  hook-error notice at every session start until someone fixes it — a check that
  cries wolf on startup is a check people learn to scroll past.
- **`set -euo pipefail` vs `set -uo pipefail` — pick by contract, and know there
  are two ways to keep an always-exit-0 contract.** A hook that may block
  (PreToolUse) wants `-e`: an unexpected failure aborting the script is
  survivable, because the caller treats a non-0/2 exit as "proceed". A hook whose
  contract is **ALWAYS exit 0** (PostToolUse injectors, SessionStart checks) has
  two honest shapes: (a) **drop `-e`** and `||`-guard every risky command —
  with `-e` on, one `grep` that legitimately finds nothing kills the hook
  mid-way and the CLI surfaces a bare `Failed with non-blocking status code`
  (pitfall #8, Pattern E's shape); or (b) **keep `-e` and add `trap 'exit 0' ERR`**
  so any failure still converts to exit 0 while `-e` keeps guarding the plumbing
  (`git-commit-headcheck`'s production shape, Pattern D). Either is correct;
  what you cannot do is `-e` alone with no trap and no `||`-guards. Rule of
  thumb: **`-e` for hooks that decide; for hooks that report, drop `-e` or trap
  it** (pitfall #8).
- **Stop is the odd one out, and the one most often reached for by mistake**:
  it's the *only* hook type that can react to what the model **itself just
  generated** (its own reply text). Every other hook type — including
  `UserPromptSubmit`, which sounds like a plausible place to police "what gets
  said" — only ever sees the **user's** input; it structurally cannot see the
  model's own *current-turn* output (that claim holds — this is still the
  right reason to route such a rule to Stop). That guarantee, however, does
  not extend to proving the `.prompt` field always originated from a
  keystroke: a background task-notification's own report text can populate
  it too, with nothing in the stdin JSON marking the difference — #30. A rule like "the model must not invent a shorthand name
  for something it hasn't verified" belongs on Stop; put it on
  `UserPromptSubmit` instead and it will (a) never once catch what it was
  built for, since that text never flows through that event, and (b)
  false-block the user's own unrelated typing whenever it happens to contain
  the trigger pattern. This is a category mistake, not a tuning problem — no
  amount of regex refinement on the wrong event fixes it. Full contract
  (`last_assistant_message` vs `transcript_path`, the anti-loop check) in
  Pattern E in [references/hook_patterns.md](references/hook_patterns.md).
- **Stop has two block channels with identical loop protections — pick by
  intent, and make the first (only) block carry everything.** `decision:
  "block"` + `reason`, or plain exit 2 + stderr, shows as a hook *error* — for
  hard gates ("this must not stand"). `hookSpecificOutput.additionalContext`
  shows as neutral "Stop hook feedback" with no error notification — for
  coaching and reminders the model should weigh, not gates. Both count toward
  the same consecutive-block ceiling from the table above, so the choice is
  tone, not safety. What that means for message design: a blocked retry
  round (`stop_hook_active: true`) is let through **with whatever violations
  remain** — so a Stop guard gets exactly **one** informed bite. (The ceiling
  reinforces this only when your remediation is "rewrite the reply"; if it
  involves tool calls the counter resets and the ceiling never lands — #27.
  Either way the one-bite conclusion holds, because it rests on the latch, not
  on the ceiling.) Report *all* findings in that
  one block (a guard that prints only the first loses the rest permanently —
  pitfall #17), and write the message as an escape manual naming the exact
  acceptable fix, not a verdict — the model converges in one round or it burns
  the cap guessing. v2.1.145+ inputs `background_tasks` / `session_crons` let a
  blocking hook tell "the session is done" from "the session is merely paused
  waiting for background work" — blocking a pause forces pointless
  continuations and wastes the same cap.

Full runnable skeletons: [references/hook_patterns.md](references/hook_patterns.md).

## The skeleton (PreToolUse Bash guard)

```bash
#!/usr/bin/env bash
set -euo pipefail
IFS= read -rd '' INPUT || true                 # builtin; NOT $(cat) — see below
# 0-fork fast path: a builtin `case` on the raw JSON, BEFORE paying for python3.
# Your guard runs on EVERY matching tool call, so the irrelevant path is the one
# that has to be cheap. Keep this filter BROADER than what you actually block and
# never flag-level — it answers "is this even about X", nothing finer (#22).
case "$INPUT" in *TRIGGER*) ;; *) exit 0 ;; esac
TOOL=$(printf '%s' "$INPUT" | python3 -c "import sys,json;print(json.load(sys.stdin).get('tool_name',''))" 2>/dev/null||echo "")
[ "$TOOL" != "Bash" ] && exit 0                # only guard the tool you mean to
CMD=$(printf '%s' "$INPUT" | python3 -c "import sys,json;print(json.load(sys.stdin).get('tool_input',{}).get('command',''))" 2>/dev/null||echo "")
[ -z "$CMD" ] && exit 0
printf '%s' "$CMD" | grep -qw 'TRIGGER' || exit 0   # precise relevance check
# ... precise detection here ...
if <command actually does the banned thing>; then
  echo "BLOCKED: ... WHY ... USE INSTEAD: ..." >&2   # stderr = the guidance shown
  exit 2
fi
exit 0
```

**Why the first two lines are not stylistic.** `INPUT=$(cat)` plus each
`printf … | python3 -c …` costs forks **on every call this hook matches, including
the ones it has nothing to say about**. A fleet of ~13 Bash-matcher hooks × parallel
sessions × sub-second tool cadence turned that into a sustained 40–200 forks/sec of
pure guard overhead and put Gatekeeper at the top of an all-day CPU ranking with no
runaway process anywhere — the fleet was fine; the *irrelevant path's* per-call cost
was the bug (#22, with the per-guard conversion recipe and its measured floor).

Three caveats before you copy the `case` line anywhere else — the first one is the
difference between a fast path and a bypass:

- **A coarse filter must be a SUPERSET of what you block, and a raw substring test
  is not one.** `TRIG''GER -x` runs `TRIGGER` — bash splices the quotes away before
  execution — but the raw event text contains no `TRIGGER` substring, so a bare
  `case "$INPUT" in *TRIGGER*)` exits 0 and the guard never sees it. **Measured**:
  drop this exact line into the shipped Pattern A and `TRIG''GER -x` flips from
  exit 2 to exit 0, a full bypass — while `scripts/test_hook.sh` still reports
  21 pass / 0 fail, because no row carries a spliced trigger. Pattern A already
  carries the fix and the reason ("de-splice — strip quotes and backslashes — and
  check again; a false negative is a full bypass"); a coarse filter placed *before*
  that de-splice makes it unreachable. Two safe shapes, in order of preference:
  **filter on something the splice cannot touch** — a JSON key or a tool name
  (`case "$INPUT" in *'"tool_name":"Bash"'*)`), since quote-splicing lives in the
  *command* text and cannot rewrite the event's own structure; or **de-splice
  inside the filter** before testing (strip `"`, `'` and `\` from a copy of the
  input, then match). Prefer the first: it needs no escaping gymnastics, and a
  filter whose own quoting you have to get right is a filter you can get wrong
  silently. The skeleton above is safe
  as written only because its own detection is likewise a plain word match; the
  moment the guard below the filter is smarter than the filter, the filter decides.

- **This skeleton is fail-open on irrelevance** (`grep -qw … || exit 0`), so a
  coarse filter in front of it changes cost, not semantics. **A fail-closed guard is
  different**: a bare substring filter silently converts its contract from
  block-unknown to allow-unknown (measured — `'not json'` sailed straight through the
  first cut of that fix), and a `*tool_name*` marker alone re-opens the same hole from
  the other side. Read #22's gate requirements *before* fitting a fast path to a guard
  that is supposed to block on malformed input.
- **There is a cheaper layer above the script.** A hook handler can carry an
  `if` field in its registration — permission-rule syntax such as `"Bash(git *)"` —
  and the hook command **does not run at all** when it doesn't match: zero forks,
  because zero processes. It is best-effort by design (the docs say it fails *open*,
  running your hook anyway, when the Bash command can't be parsed), so treat it as a
  cost optimization and **never as the gate** — the in-script check still decides.
  Three sharp edges: it holds exactly one rule (no `&&`/`||`), it is only evaluated
  on tool events, and a hook that sets `if` on a non-tool event **never runs at all**.

## Rules that separate a working guard from a session-poisoning one

Not style preferences — each is a specific failure we shipped and traced back.

### 1. Match at the **token level with shlex**, never awk-split the raw string

A guard that **false-blocks a healthy command is worse than one that misses** —
a guard people must bypass gets bypassed reflexively, and then it protects
nothing (the core discipline: *误杀健康输入比漏报更糟*). The recurring cause of
false-blocks is matching on the raw command string.


- **Wrong**: `awk '{gsub(/&&|\|\||;|\|/,"\n")}'` to split into segments — awk
  doesn't understand shell quoting, so `grep -E "a|TRIGGER|b"` gets split at the
  `|` *inside the quoted regex*, `TRIGGER` becomes a phantom command, and the
  guard blocks a plain grep. (Shipped 2026-07-21; the guard's very first real use
  was a false-block on my own grep.)
- **Right**: tokenize the whole command with the **`shlex.shlex` class**, not the
  `shlex.split()` function — `split()` only treats `| ; & < >` as separators when
  they are space-separated, so `ls|TRIGGER x` tokenizes to `['ls|TRIGGER', 'x']` and
  your command-position check never sees `TRIGGER` at all (measured; the class with
  `punctuation_chars=True` yields `['ls', '|', 'TRIGGER', 'x']`). Copy a shipped
  walker verbatim rather than reaching for the one-liner — but **copy the one that
  passes `scripts/test_hook.sh`**, which is **Pattern A's**. The
  [walker section](references/hook_patterns.md#the-shlex-command-position-walker) is
  a *compact* form and says so: it omits the per-wrapper valued-flag tables, so it
  misses a target riding a **valued-flag wrapper** — measured, it returns "not in
  command position" for `timeout 5 TRIGGER`, `sudo -u root TRIGGER` and
  `nice -n 10 TRIGGER`, while Pattern A's version catches all three. (Bare
  `sudo TRIGGER` is fine in both — it is the wrapper's *own* flag taking an argument
  that the compact table doesn't know to skip.) Only one of those shapes is in the
  shipped harness, so the run you actually see is **20 pass / 1 fail on
  `wrapper-timeout`** against 21/0 for Pattern A's; the other two fail silently
  because no row covers them. A quoted
  `"a|TRIGGER|b"` stays **one token**, so a regex argument is never mistaken for
  a command. Then check whether your target is in a **command position**
  (token[0], or right after a `;`/`&&`/`||`/`|` separator, skipping `VAR=val`
  env-assignment prefixes). Command-position walker in
  [references/hook_patterns.md](references/hook_patterns.md).
- Corollary: `echo "…TRIGGER…"`, `grep TRIGGER`, `# TRIGGER`, `man TRIGGER` must
  all pass. Your test set MUST include these mention-not-execute cases.
- **Corollary — exempt `git` write segments before they reach the walker.** A
  commit message is arbitrary data, and the whole message text reaches your
  command-position walk as pseudo-command-text — `git commit -F - <<EOF` with a
  body quoting `foo|TRIGGER` lands `TRIGGER` in command position, and the guard
  blocks its own fix commit (pitfall #7 is exactly this, shipped). Any Bash guard
  that inspects command strings must skip segments whose head is `git` +
  `commit`/`rebase`/`tag`/`am`/`cherry-pick` — and do it at the whole-command
  level, before any line splitting (Pattern A shows the order; the production
  version is `lib-git-commit-detect`'s adjacency check).
- **Corollary — the walker is two-stage for a reason.** `whitespace_split=True`
  treats newlines as ordinary whitespace, so a multiline block
  (`cd /x\ngit add\nTRIGGER -y`) collapses into one segment headed by `cd` and
  the trigger is never in command position — replayed trigger rate 0 on real
  transcripts (pitfall #11). Split into lines **shell-aware** first (quote state
  and backslash continuations honored, so quoted multiline strings don't
  fragment), then shlex-walk each line — both Pattern A and the walker section
  ship that splitter (`split_shell_lines`, production-proven in qlmanage-guard).
  What even it cannot parse is a heredoc body (not quote syntax); when to accept
  that residual is #11's call.
- **But shlex isn't a silver bullet, and *what* you detect changes whether
  fail-open is safe.** `shlex.split()` itself throws `ValueError` on an unbalanced
  quote — a multi-line `git commit -m "…` message with a `#` or an unclosed quote
  is the classic trigger. The `except ValueError: cmd.split()` fallback then
  *allows*, which is right when you're detecting a **banned modifier** (does this
  carry `--no-verify`? — missing it errs safe, Rule 1's direction), but
  **dangerous when you're detecting whether the command IS your target at all**
  (is this a `git commit`? — a ValueError there means the guard never recognises
  the commit and silently doesn't fire; a real cross-domain commit shipped with no
  confirmation dialog this way). For the *is-this-the-command* decision, prefer a
  narrow **regex** (`git` and `commit` as separate words, any flag tokens between)
  that's immune to multi-line-quote breakage; reserve the shlex walker for the
  *command-position / modifier* checks where fail-open is the safe direction.
  (The boundary: regex when the predicate is "is this a specific common command
  at all" — `git commit`, `git push` — whose own message/arguments are what breaks
  tokenizing; walker when the predicate is "is a *banned* command or modifier in
  command position" — there the banned thing is rare and a ValueError fail-open
  errs safe, Rule 1's direction.)

### 2. Test with **bash -n + a real JSON event, end-to-end, BEFORE registering**

**A corrupted or wrong-logic PreToolUse hook poisons the *entire* session** —
every later Bash call gets truncated / duplicated / falsely-failed / looks
hallucinated-executed, and you'll blame "the environment" when it's the hook you
just installed. (2026-07-05: a `[^;&|]` regex broke in one edit, `;&` became a
bash case-fallthrough token, poisoned half a session until `bash -n` found it.)
"My tests passed at deploy" isn't enough — the file can corrupt in a *later* edit.

Gate before registering ANY hook:
```bash
bash -n hook.sh                                # syntax
printf '%s' '{"tool_name":"Bash","tool_input":{"command":"<trigger case>"}}'    | ./hook.sh; echo "exit=$?"  # want 2
printf '%s' '{"tool_name":"Bash","tool_input":{"command":"<healthy lookalike>"}}'| ./hook.sh; echo "exit=$?"  # want 0
```
Bundle the harness: [scripts/test_hook.sh](scripts/test_hook.sh) runs a whole
table of trigger/allow cases. **Self-block gotcha:** once the hook is live in the
session you cannot test it by putting the trigger string in your *own* Bash
command — the live hook blocks your test command. Put the cases in a **script
file** and run `bash test_hook.sh`; the outer command doesn't contain the
trigger, so it isn't self-blocked.

**Once a hook has caused one real incident (a false-block or a silent miss),
solo re-reading the code is not enough** — a same-day rewrite of a Stop-hook
guard was itself re-broken twice by the author while fixing the first bug (a
quote inside a Python comment, invisible on re-read, only surfaced by running
the actual failing JSON case). The escalation is a multi-lens agent-team
review where every finding must be reproduced by *executing* a real payload
against the live script, not by reading the code and agreeing — this is the
general Counter Review methodology
(skill-creator's `skill-development-methodology` reference, Phase 6), applied to
a hook instead of a skill. In one such pass, 3 lenses (matching
logic / shell-embedding safety / event-contract robustness) surfaced 13
confirmed, independently-reproduced bugs and 1 finding whose own cited
evidence turned out to be a hallucinated doc quote — caught only because the
verifier was required to curl the raw source and grep for the exact string
rather than trust the citation.

### 3. SSOT + symlink so a reinstall can't silently disarm the guard

Real script in a version-controlled dir, **symlinked** into the hooks dir Claude reads:
```
~/scripts/claude-hooks/<name>.sh      # SSOT (this setup: a private git repo)
~/.claude/hooks/<name>.sh             # symlink → SSOT

# install / recover:
ln -s ~/scripts/claude-hooks/<name>.sh ~/.claude/hooks/<name>.sh
```
A `~/.claude` reinstall wipes the hooks dir; the symlink target survives, and
recovery is one `ln -s`. A dangling symlink disables a Tier-0 guard with **zero
signal** — which is why a SessionStart health check exists (rule 4; runnable
skeleton: Pattern C in [references/hook_patterns.md](references/hook_patterns.md)).

### 4. Registration is per-profile — converge ALL profiles, release via a HUMAN gate

- **There are two registration surfaces, and the health check probably only
  watches one.** Besides the profile settings below, a repo's own
  `.claude/settings.json` (and `settings.local.json`) registers hooks for
  sessions in that repo. Those live outside `~/.claude/hooks/`, so a guard-rail
  check that walks that directory covers none of them — syntax, path,
  `--selftest`, nothing. Two silent failures shipped in one such file and
  survived two months under a SessionStart health check built to prevent
  exactly that (#38, #39). Register project hooks by **absolute** path — a
  relative one resolves against the session cwd and breaks the first time
  someone starts `claude` from a subdirectory — and extend the health check to
  walk up from the event's cwd for project settings.
- A hook in `~/.claude/hooks/` protects **nothing** if the *active profile's*
  `settings.json` doesn't call it. Multi-profile users ran with zero guards until
  every profile was converged. Register in the **main** profile's settings
  (`~/.claude/settings.json` in this setup; the Registration section of
  [references/hook_patterns.md](references/hook_patterns.md) has the exact jsonc
  shape) — PreToolUse → matcher `Bash` → your hook — then converge the rest (this setup
  uses `sync-profile-settings.py --all`, owned by the `claude-switch-models-setup`
  skill). A SessionStart health check greps each profile for the Tier-0 guards to
  catch drift. Settings edits are picked up by the CLI's file watcher (official
  hooks docs), so registration is live without a restart — confirm by watching
  the guard fire on a safe probe, or at next session's health-check line.
- **If the guard needs a release valve, make it a *human* gate, not an env var.**
  A static `GUARD_OK=1` escape hatch is no gate — the model can set the env var
  itself. Use a native macOS dialog (`osascript` — the model can't click);
  refuse/cancel/timeout = hard NO; log every prompt/bypass to an audit file.
  Pattern in [references/hook_patterns.md](references/hook_patterns.md).
  - ⚠️ **`/dev/tty` is not a second channel — the docs say hooks cannot open it.**
    This file used to prescribe a typed `YES` on `/dev/tty` alongside the dialog.
    The official reference is explicit: hooks "run in their own session **without a
    controlling terminal**", and "the hook process and any child processes **can't
    open `/dev/tty`**" (`terminalSequence` is the documented replacement for writing
    to it). So a "two-channel" gate built that way is one channel plus dead code,
    and on a box with no GUI session the gate can never be approved by anyone.
    Consistent with local observation, though read the boundary carefully: in one
    setup's shared audit log — 1,801 entries, several guards writing to it, three of
    which implement a tty channel — **360 lines carry a channel tag (236 dialog
    confirmations, 124 declines or timeouts) and not one line of any kind names the
    tty channel.** That means the tty branch was never *entered*, which on macOS is
    what you would predict anyway, because the dialog answers first and short-
    circuits it. So the log shows nothing here ever depended on tty; it is the
    documentation, not this measurement, that establishes tty cannot work at all.
    Keep the dialog; if you need a non-macOS gate,
    you need a channel this file does not yet have a verified answer for.
  - ⚠️ **A human gate that outlives the hook timeout fails OPEN.** Hook `command`
    timeout defaults to 600s (30s on `UserPromptSubmit`), and a timed-out hook
    **does not block the tool call** — so an unanswered dialog does not become a
    "no", it becomes an allow. Bound your wait well under the timeout and make
    no-answer resolve to block *yourself*, before the harness resolves it for you.
  - The docs also carry an in-UI channel — PreToolUse `hookSpecificOutput`
    `permissionDecision: "ask"`, which prompts through Claude Code's own interface.
    It is worth knowing about, but **unverified here under `bypassPermissions` /
    auto-accept**, which is precisely the mode a Tier-0 gate must survive; the
    dialog is prescribed because it does not depend on permission mode.
  - **Below Tier-0, where a model-serviceable escape *is* allowed, make it the
    correct usage rather than a bypass flag.** The rule above is absolute for Tier-0
    and does not bend here — this is about the correctness guards that fall short of
    it, which still need a way out for the legitimate case the detector cannot
    distinguish. The question is what you make that way out *be*. A `SKIP=1` /
    `--force` env escape trains exactly the reflex rule 1 warns about, and under
    deadline it is indistinguishable from a bypass. Prefer an escape that is **the
    thing you wanted them to do anyway**, so taking it improves the command instead of
    disarming the guard: `pipe-fallback-guard` exits 0 the moment the command mentions
    `pipefail` / `PIPESTATUS` / `pipestatus`, because an author who wrote any of those
    has already demonstrated they understand pipeline exit codes — the guard has
    nothing left to teach them. The test to apply: *if someone takes my escape hatch,
    is the resulting command better, or merely unblocked?* If the honest answer is
    "merely unblocked", you have a bypass flag with a nicer name. Sibling principle
    for the Tier-0 case, where the override exists only so the gate can be tested:
    **the escape hatch may only make the gate stricter** — see the `GIT_GUARD_TEST`
    discussion in [references/hook_patterns.md](references/hook_patterns.md).

### 5. Decide the failure **direction**, and test *that* — not just the happy path

Rule 1 ranked *detection-tuning* errors: given that the guard ran, false-blocking a
healthy command beats missing a rare bad one, because a guard people must bypass
gets bypassed reflexively. **This rule is about a different axis — the guard's
machinery not running at all** — so "which is worse" is not being reversed here;
the two rankings never meet. A tuning miss costs you one case; this costs you the
guard, silently, on every input of that shape.

The failure: the guard **cannot obtain the thing it judges on** — a parse throws, a path doesn't
resolve, a dependency is missing, a subprocess times out — and the very
`2>/dev/null || true` that stops the hook from crashing quietly converts *"I could
not check"* into *"nothing to report."* The hook exits 0. **That output is identical
to a real pass**, which is why this survives for weeks.

So at every point where the hook *obtains* something (parses the command, reads
staged files, queries a service), decide explicitly: **if this comes back empty, does
that mean allow or block?** — and write the answer next to the branch. Fail-open is
often right for a *modifier* check (does this carry `--no-verify`? missing it costs
you one case). Fail-closed is usually right for the *is-this-even-the-thing* check
(is this a cross-domain commit? an empty answer means the guard never fired at all).

**Then test the direction, not the happy path**: hand it an unresolvable path or an
unparseable command on purpose and assert it still does what you decided. A suite
where every row passes *because the hook silently allowed everything* is
indistinguishable from a suite that passes.

**Read those results carefully — the same input has opposite correct answers for
different guard classes.** Take `cd ~/no-such-dir && TRIGGER`:

| Guard class | Judges on | Correct exit | Why |
|---|---|---|---|
| **Token matcher** (is this a banned command form?) | the command text alone | **2, block** | `TRIGGER` is right there in the text; an unresolvable `cd` doesn't make it not-a-trigger, and if the guard goes quiet here it will also go quiet on `cd ~/real-dir && TRIGGER` |
| **State deriver** (does the repo's staged set span domains?) | state read from disk | **0, allow** | `cd` fails, `&&` short-circuits, no commit ever happens — there is nothing to guard |
| **Termination-state reader** (has the remediation already happened?) | a receipt / counter file (rule 7) | **0, allow** — *when the state file IS the termination condition* | an unreadable receipt means the hook cannot know it already fired; failing closed here blocks forever with no remediation possible and no human-visible cause — that *is* the loop, and it is the one failure worse than a missed case. **Inverted sub-case — read this before copying the row:** when the state is only a **budget on top of an independent predicate** (the block still clears by doing the work), allow-on-unreadable **silently disables the entire hook** — one unwritable directory makes it mute for every input, forever, which is the worst failure shape there is. There, fail back to *the behavior before the budget existed* (keep evaluating the predicate), not to silence. **Tell the two apart with one question: if the state vanished, would remediation still be possible?** No → receipt case, allow. Yes → budget case, keep checking. Worked answers, so nobody has to re-derive them: rule 7's mechanism 2 (receipt) **and** mechanism 3 (per-session counter) are both **receipt case → allow** — mechanism 3 is deliberately blind to whether R happened, so its counter is the only exit and muting it strands the turn. The budget case is a counter layered on a predicate the user can still satisfy on its own |

So decide which class your hook is *before* writing the row, and the harness's
`unresolvable path` template row expects **2** because that template targets the
token-matcher class. Getting this backwards produces a confident FAIL against a
correct guard. For a state-deriving guard the failure you are hunting is: **the command would really
have run and the guard didn't see it** — an unbalanced quote makes tokenizing throw,
the fallback allows, and a genuine cross-domain commit ships with no dialog (rule 1's
ValueError note). Ask of every allowed row: *would this command actually have done
the thing?* If no, the allow is correct.

Running this exact probe against a real state-deriving guard returned two allows on
the first pass: one was correct (the short-circuit above) and one was a genuine
fail-open. **The probe finds things; you still have to classify what it found** —
which is why the class table above comes before the rows.

Real case (2026-07-22): a scope guard read staged files via `git -C "$REPO_DIR"`
with `REPO_DIR` parsed out of the **command text** — so `cd ~/repo && git commit`
handed it a literal `~/repo`, `git -C` failed, staged came back empty, and the guard
concluded "no cross-domain files, allow." Every cross-repo commit went unguarded and
nothing ever looked wrong. Anatomy + the shared-library twist: pitfall #10.

That parser has a second failure direction, and it is the nastier one. Once you
add a fallback so it stops failing open, the fallback becomes correct for one
reason and wrong for another — and both print the same line. `git push` (no
explicit target) legitimately falls back to the event's `cwd`; `git -C "$R" push`
*names* a target the hook cannot resolve, falls back to the same `cwd`, and then
renders a confident ✅ about a different repository. Those two cases render
byte-identically (measured, MD5-equal), so neither the hook nor the reader can
tell the honest verdict from the misbound one. **A fallback value must carry the reason it was
chosen**, and only "no explicit target" earns a verdict. Full anatomy, the
confused-deputy framing, and why fixtures with literal paths never catch it:
pitfall #28.

### 6. Judge on a fact the world can answer — never on your own rendering, never on a naming habit

Rules 1 and 5 are about *how* you match and *which way* you fail. This one is
about **where the thing you match on came from**, and it has two failure shapes
that both go silent:

- **Never branch on a string you formatted for a human.** If the hook builds a
  report — sorted, joined, truncated to the first N with a `(+M more)` tail — and
  then pattern-matches its own decision against that report, the branch inherits
  the rendering's losses. Items past the cutoff simply do not exist to it, so the
  branch works on every small fixture and stops firing on exactly the large
  sessions it was built for. Emit the machine fact on its own channel (one
  untruncated `KINDS:a,b,c` line) and match *that*. A rendering is an output, not
  a data source (pitfall #12).
- **Prefer a checkable fact over a naming convention.** Classifying by path shape
  (`/skills?/[^/]+/references/`) encodes one directory layout; a repo laid out any
  other way is classified `None` — silently, forever. The fix is *not* to widen the
  pattern, which trades a silent miss for machine-wide false positives (rule 1
  forbids exactly that trade); it is to ask a question the filesystem can answer —
  *is there a `SKILL.md` beside this `references/` directory?* Facts survive
  layout changes; conventions do not. (When the candidate **is** a `SKILL.md`,
  there is no sibling to ask about — classify by basename; #13 explains why that
  is a spec-defined fact and not the naming habit this rule warns against.)
  **A checkable fact can still be the wrong fact — anchor the question and filter by
  type.** `test -f SKILL.md` is true in a downloads folder too, and one guard that walked
  ancestors looking for exactly that swallowed an entire home directory, then told a real
  session to load a skill named after it — a name that cannot exist (rule 9's incident).
  Anchor to a sibling of a *specific* directory or to a known install path; an unanchored
  ancestor walk is a convention wearing a fact's clothes.

The tell for both: a branch that has never once fired in production while its
tests are green. Print the raw pre-formatting classification and you will see
which of the two you have.

### 7. If the hook **demands remediation**, prove the loop terminates

A hook that **blocks** (exit 2) until X is done — Stop hooks especially, since
they re-fire on every subsequent stop — is not a check, it's a **feedback loop**.
(A hook that merely *injects* a demand and exits 0 has no loop at all: nothing
re-evaluates. That is mechanism 0 below, and it is the right default more often
than people reach for it.)

```
condition T is true → hook demands remediation R → model performs R → T checked again
```

**Write the Loop Contract before the first cycle — for hook-enforced loops and
agent-driven review / wait / retry loops alike:**

```text
LOOP KEY: immutable logical target / lineage + one failure axis
FIRE T: the condition that starts another cycle
REMEDIATION R: the exact action one cycle performs
VARIANT V: the well-founded quantity that strictly decreases for this key
BUDGET: maximum cycles, fixed before cycle 1
SUCCESS EXIT: the observable that proves the axis is clear
CAPPED EXIT: what is left blocked / unshipped / pending when the budget ends
```

No completed contract means no blocking Stop hook and no repeated reviewer or
polling loop. Freeze the key before cycle 1. A remediation snapshot, commit, or
reviewer name stays inside that same lineage and cannot mint a new budget. A
new, unrelated finding is a **new key**: record it separately; it does not reset
this loop's budget. A cycle that cannot name a new falsifying experiment or a
smaller V adds no evidence and stops.

For an **agent-driven independent-review loop**, the default budget is one
initial review plus one narrowly scoped re-review after substantive fixes. A
third reviewer is not automatic. If the re-review still reproduces a BLOCKER or
MAJOR on the same axis, leave the hook unregistered / artifact unshipped, report
the blocked state, and require a new user-authorized task whose Loop Contract
declares its budget before cycle 1. An agent-declared budget cannot authorize
itself. Inside that authorized task, name the concrete safety or business
failure caused by stopping now; optional polish does not qualify.

Filled review-loop example:

```text
LOOP KEY: <initial frozen commit>'s review lineage + termination-contract fidelity
FIRE T: fresh review reports a same-axis BLOCKER / MAJOR
REMEDIATION R: reproduce that finding, apply one bounded fix, run its narrow check
VARIANT V: 2 - completed review cycles
BUDGET: 2 cycles total (initial review + one re-review)
SUCCESS EXIT: no same-axis BLOCKER / MAJOR
CAPPED EXIT: artifact stays unregistered / unshipped; report remaining findings
```

Every repair descendant of the initial frozen commit remains in this key. The
current snapshot changes so the reviewer can inspect the fix; the lineage and
its remaining budget do not.

Nothing mechanically enforces this hookless budget — it holds only while the
agent follows the Skill. That limitation is why the capped exit must be visible
and must never be reported as “completed.”

**If completing R can make T true again, the loop does not converge.** Nothing
errors, nothing crashes; it burns round after round until a human interrupts —
which is what usually happens, because each round is a *complete* remediation
cycle (dispatch, wait, adopt, edit

…(truncated)
