Reality Audit — spend the tail on the thing that ships
Implementation evidence is easy to collect from the wrong surface: a scratch build, a one-off interpreter run, a test that bypasses the public boundary, plausible printed output. This terminal protocol decides whether the behavior that ships has actually been exercised.
Load it after the final material implementation change, before saying verified, done,
green, or no work remains — and again after an owner correction or a reproduced falsifier
invalidates earlier evidence. Not at task entry, and never interrupt implementation with graph
modelling to prepare for it.
"Checked" is a two-step claim, and the step is always declared. Checking against the record
— look / orient / trace, whether the words agree with what is written — is one step; the
graph is an instrument for observing the world, not the world. Witnessing the world — fresh
observation on the canonical carrier — is the other. Record-checking ranks below world-witnessing,
and behavioral closure is carried by the second only. So when you close anything with the word
"verified", say what you verified it against; a verdict with no named step is not a result. Every
verdict this protocol issues is a world-step verdict.
Proportionality. When the change alters no public contract — no new or changed exported symbol, output format, config/schema shape, or documented behavior — the ladder collapses to steps 1–2 on the touched path: rebuild the canonical deliverable, exercise the changed path once with an asserted expectation, record command and exit status in one line. The full claims table is priced for behavior-changing work, not typo-class fixes. In doubt whether a contract changed? It did.
1. Freeze only the load-bearing claims
Read the accepted source, not the implementor's completion summary. Contract precedence:
- latest owner correction;
- accepted requirement or specification;
- established public API, schema, and canonical tests in the target version;
- implementor relay — evidence about fulfillment, never the contract.
Record only claims whose failure changes acceptance. Exact output, ordering, error, serialization, or command examples are executable claims even when written as prose. Reproduce each normative example exactly; a simpler representative case does not discharge it.
2. Name the surface and the falsifier before running them
For each claim, one compact row:
Take the surface from the config, don't re-derive it. A verstakified repo answers this in
AGENTS.md → Reality: the canonical carrier per claim class, the exact observation, and who can
reach it. Read that table first — the carrier is the owner's to name, not the auditor's to guess.
No table, or a class it doesn't cover: ask the owner and offer a verstakify run, and record
what they answer so the next audit starts from it. Deriving a carrier yourself is the guess this
whole protocol exists to remove. Its Ceiling rows are binding: a claim class recorded there as
having no reachable observation cannot come back verified from this protocol, however the probe
goes — provisional or blocked is its honest top.
| Field | Required answer |
|---|---|
| Claim | What exactly must hold? |
| Canonical surface | Which exact binary, package, exported symbol, API route, UI, schema, generated artifact, or deployed path will users consume? (From Reality where the repo has one.) |
| Observable | What should an external observer see? |
| Executable falsifier | Which command, request, assertion, or test would prove the claim false? |
Boundary words are equivalence classes, not one convenient fixture. For empty, none,
missing, malformed, default — enumerate the distinct representations the public parser can
receive (zero bytes vs a blank record, an empty array vs a missing field, an omitted flag vs its
explicit default) and probe every representation the accepted wording equates. For a minimum,
maximum, threshold, window, count, or version gate — the accepted boundary plus the nearest value
on each side with different behavior; a default that supplies the boundary is tested both omitted
and explicit. A parser, validator, or helper reused from a stricter sibling path gets its
preconditions checked independently: an inherited guard must not reject input the new public
contract accepts.
Name the canonical path before the probe. A scratch build, local script, mock-only call, or test
that bypasses the public boundary is provisional unless it is the exact public deliverable.
Authorship does not decide: a fresh black-box test verifies when it rebuilds and executes the
canonical surface with the exact falsifier; a same-author mock or internal-only test does not.
When several artifacts could be invoked, prove which one runs — path plus freshness (mtime,
digest, version, or build marker). Printed output is observation, not a pass, until an executable
assertion checks the expected value and exit status.
3. Run the terminal evidence ladder
After the final material change, in this order:
- Build/package/deploy the canonical deliverable — before testing any scratch copy.
- Exercise the newly changed path at its public boundary. Assert the expected observable and failure behavior; do not merely inspect plausible output.
- Re-run the exact prior falsifier. A contradicted claim stays contradicted until the same case, or a strictly stronger one, passes against the fresh deliverable.
- Re-run exposed old requirements in code. Acceptance cases touching changed shared code first; a broad suite only after the exact new path has run.
Any material artifact change after step 1 invalidates the tail: rebuild and repeat the affected probes. Never let cleanup, graph work, or narration consume the budget of the new-path probe.
NKS comes after the evidence and is never part of it. Freeze the claim verdicts and finish all
material artifact changes first; then at most one terminal update, and only when a durable
correction or contradiction will change a later agent's decision. A later material patch makes
that update premature — its behavioral confidence is provisional until the affected public
evidence is rerun and the node reverified. The graph is not mid-audit scratch, and no graph state
raises a behavioral verdict. What is forbidden is the substitution, not the reading: the
record-step is legitimate and obligatory in its own place — resolving the contract, deduping,
integrity's claim-audit — and only passing it off as world-witnessing, or letting it lift a
behavioral verdict, is illegitimate. The one-update limit scopes this audit's own evidence handoff; the
structural modelling writing or the repo's push ritual requires belongs to the round, not to
the audit.
Exit status is the verdict (the canonical statement — writing points here). Combined
terminal evidence fails closed from its first command (set -euo pipefail or the platform
equivalent); an expected failure code is captured and asserted locally, then fail-closed execution
is restored. A claim is not verified while its evidence command exits nonzero: green subtests,
passed lines in stdout, or a later unrelated zero do not override the recorded exit code — read
it from the tool result, never infer it from selected output; only a rerun of the same case (or a
strictly stronger one) against the fresh deliverable clears it. A wrapper that deliberately runs
every subgroup accumulates failures and exits nonzero when any required subgroup fails. Before the
final narration, reconcile every passed/green claim with the recorded exit statuses and the
failing subgroup summary.
4. Give one truthful verdict per claim
- verified — the fresh canonical surface produced the required observable and the executable falsifier was attempted successfully;
- provisional — evidence exists but is scratch-only, mock/internal-only, print-only, stale, or misses the exact public boundary;
- contradicted — a reproduced counterexample still fails;
- blocked — the evidence surface is unavailable; quote the literal blocker and name only the claim it prevents checking.
Report a compact claims × verdicts table with the canonical path and command/test evidence — actual telemetry only, never estimated tokens, cost, duration, or coverage. A green carry-over subset that omits the failing case is absence of observation, not repair. Structural graph health is a separate fact and never upgrades a behavioral verdict.
Required work closes only when every required claim is verified or the owner consciously accepts a named exception. Provisional, contradicted, and blocked are handoff states, not synonyms for done.