Agent release gate
Product-level sanity QA for the agent runtime, one layer below the playground UI. The question
is not "is every detail right" — it is "if a user opens the product and does the obvious first
things, do they work?" This is the gate a release passes before shipping.
Every check asserts on the wire (the SSE frame types the browser sees) and on side effects
(the file really persisted, the revision really incremented) — never on what the model says. That
makes it deployment-agnostic: point it at any stack and the assertions still hold.
Run it
Set three environment variables for the deployment under test, then run the gate:
export AGENTA_BASE=https://your-stack.example.com # deployment origin
export AGENTA_PROJECT_ID=... # target project
export AGENTA_API_KEY=... # project API key
uv run resources/qa_product.py --all --custom-slug <vault-slug> --custom-name "<display-name>" --require-store # everything
uv run resources/qa_product.py --cell P1 # one cell
uv run resources/qa_product.py --cell C1 --only chat # one journey
uv run resources/qa_product.py --cell S2 --only warm --only cold1 --require-store # continuity
EXPORT the three variables, do not just set them. The driver falls back to an env FILE when
AGENTA_* is absent from its environment, which is helpful interactively and dangerous in a
release run: a credentials file of bare KEY=value lines sourced with . file sets the shell
only, the child uv run process inherits nothing, and the driver silently runs the whole gate
against WHATEVER DEPLOYMENT the fallback file names. The failure surfaces as 401 Invalid credentials from a stage whose key you just watched answer 200, or worse as a green run
recorded against the wrong stack. Use set -a around the source, or export each variable,
and confirm the stage in the results before trusting them. (Cost a staging gate run on
2026-08-28; the fallback degrades to "wrong deployment", never to "no credentials".)
Paths are relative to this skill's directory. The deployment's vault must hold the provider keys
the cells use (Anthropic / OpenAI / OpenRouter). If the three env vars are unset the driver stops
immediately and names exactly what is missing; a legacy --env-file <path> fallback also exists.
--all includes cells P2, P2b, and P3 (a custom OpenAI-compatible provider, with P2 and P2b
running locally and P3 on Daytona), which need a vault slug passed via --custom-slug; the driver fails
fast if it's missing. --custom-name is also required because model_keys is built from the
display name rather than the stable slug. Custom-provider cells skip the credential-rotation
journey because their value is write-only and cannot be safely restored. Cells S1, S2 and C1
additionally need
the subscription sidecar logged in on the target deployment — see resources/coverage.md for what
each cell requires. The Daytona cells (C2, C4, P3, X2) additionally run the secret_opaque
journey and need the runner's Daytona API key to manage Secrets, because credential hiding is on
by default; without it those cells fail at sandbox creation with an error naming the permission.
The one flag a release conductor must not skip past. The continuity journeys
(warm, cold1, cold2) only mean anything on a store-backed deployment: with no object
store the runner degrades silently to an ephemeral working directory, so those journeys SKIP by
default and FAIL with --require-store. Run the gate against a deployment with AGENTA_STORE_*
configured and pass --require-store, or the greenest possible run still says nothing about
durability. cold2 additionally needs an operator hook that SIGKILLs the runner replica
(--cold2-replace-cmd) and SKIPs without it.
The flag that makes the gate fit the release: --release-base. The matrix is fixed, so
without it a release that reworked a subsystem gets exactly the coverage of a release that did
not touch it. Pass the ref the release branches from, and the driver reads the release's own
changed paths, matches them against the rules in resources/path_triggers.py, and makes the
cells those rules name MANDATORY for this run:
uv run resources/qa_product.py --all --release-base origin/main --require-store # every release run
uv run resources/path_triggers.py --release-base origin/main # preview only, runs nothing
A mandatory cell that lives in qa_product.py is added to the run even when --cell did not ask
for it. A mandatory cell that is a standalone matrix_*.py script is a separate process the
driver cannot observe, so it is printed, written to mandatory.json, and listed in summary.md
under "Mandatory for this release" — the release is not green until each of those has a recorded
result of its own. If a rule names a cell that does not exist, the driver stops before running
anything and says so: the release changed code a rule protects and the coverage was never
written, which is the one outcome that must never read as green. Add --changed-path to state
paths by hand where the checkout is not the release branch. With no --release-base and no
--changed-path nothing changes, so every existing invocation behaves exactly as before.
Adding a rule is one line in PATH_TRIGGERS (a glob, and the cells it makes mandatory) plus the
cell it names. Matching is fnmatch over the whole repo-relative path, so * crosses directory
separators and a/b/* covers the whole subtree; write ** so a subtree rule reads as one.
Reading the result. Each journey prints PASS, FAIL, or SKIP with a one-line reason, and
a per-cell markdown table lands with the full JSON in ./qa-gate-runs/<timestamp>/ (override the
location with AGENTA_QA_RUNS_DIR). Runs are written to the current working directory, never into
the skill. SKIP is expected where a journey does not apply to a cell (for example mcp on any Pi
cell — user MCP is Claude-only). Any FAIL blocks the release until triaged.
A SKIPped integration test in a security- or concurrency-bearing area is a FAILURE in your
summary line, not green. State it as "N passed, M skipped OF WHICH k are untested claims" and
name the k. A commit-lock race test skipping for want of a reachable Postgres is exactly how a
one-line syntax error (SET LOCAL lock_timeout with a bind parameter, which Postgres rejects
outright) survived 1911 green tests before a human hit it as his first live action.
The two journeys that run many things at once: burst and crosstalk. Every other journey
drives one run at a time, so the gate only ever saw faults that reproduce on a quiet deployment.
The credential-delivery fault of AGE-4249 does not: about one production first message in five
failed because some fresh Daytona sandboxes start without their Secret substitution wiring, and
per cold sandbox that is roughly an 8 percent fault. burst sends 16 first messages at the same
time on 16 brand new sessions, so the run buys 16 cold starts instead of one. crosstalk runs 3
two-turn conversations with long output beside 2 approval flows, and checks that no stream carries
another session's nonce, except on the codex harness, where the gate rides a platform tool with
empty arguments, so the approval command carries no nonce and isolation is not checked there
(nonce_checked=false). Both are Daytona-only by default and skip elsewhere; both report the
runner's stable error code per run, so a credential_delivery_failed names itself.
uv run resources/qa_product.py --cell C4 --only burst --only crosstalk
uv run resources/qa_product.py --cell C4 --only burst --burst-size 24 # more cold starts
uv run resources/qa_product.py --cell C3 --only crosstalk --concurrency-everywhere # local too
uv run resources/test_qa_product_concurrency.py # offline tests
Read a green burst honestly. At an 8 percent per-cold-start fault rate, 8 runs miss the fault
51 percent of the time, 16 miss it 26 percent, and two Daytona cells at 16 miss it about 7 percent.
A PASS is a sample, not an all-clear, and the result says so in its own reason line. A FAIL is
proof.
Each concurrent run holds its own Daytona sandbox, about 5 GiB of the organization's disk, and a
parked sandbox keeps counting until its auto-delete window closes, so a burst of 16 is about 80
GiB in flight. The counts are --burst-size (default 16), --crosstalk-conversations (default 3)
and --crosstalk-approvals (default 2). The cap is 32 concurrent runs: 32 for the burst size, and
32 for the two crosstalk counts TOGETHER, because what costs disk is what runs at once. When the
provider refuses on capacity the journey reports SKIP with a loud reason, never a PASS or a FAIL,
because nothing about the product was measured. --concurrency-timeout (default 300s) bounds one
TURN and rides into the stream as an absolute deadline, so a two-turn crosstalk run gets twice
that and a stream that never ends is abandoned rather than followed.
A release that changes services/runner/src/engines/sandbox_agent/** or
services/runner/src/providers/daytona* makes the Daytona cells C2, C4 and X2 mandatory through
path_triggers.py, and forces burst and crosstalk into the run even when --only named
something else. That is how these journeys reach a release that needs them.
Before a human gets a deployment URL, run resources/qa_commit_approval.py too. It is not
part of qa_product.py's cell × journey matrix — none of that matrix's journeys drive a live turn
against a REAL, saved workflow revision (the commit journey only exercises the REST API; chat,
tool, approve/deny all run against an inline, unsaved config), so nothing else in the gate
observes the S3b single-use execution-authorization gate actually firing around a real config
mutation. This script does: create a real workflow + revision, invoke a live agent turn that calls
read_config then commit_revision, expect the pause, approve it in-band, and verify the new
revision landed with REST fetch-back. Treat a FAIL here as blocking, the same as any other gate
FAIL.
Tiers: coached vs. mechanism-blind
Every cell in this gate (qa_product.py's journeys, qa_commit_approval.py,
qa_probe.py, and all matrix_w*.py cells) declares which of two tiers it belongs to, in its
own docstring:
- Coached (backend-path test). The prompt names the mechanism verbatim — which tool, which
operation, which target path. This proves OUR CODE works (the gate, the base check, the
authorization handoff, the sandbox path) when the right call is made. It proves NOTHING about
whether a model finds that call from a plain-language human ask.
- Mechanism-blind (model-behavior test). The prompt is phrased the way a real user types —
no tool names, no operation names, no schema hints. Only these cells license a claim about
what the model can do unprompted.
Rule: claims of model behavior may only cite mechanism-blind cells. Most cells in this
directory are coached tier — they test backend paths, correctly, but do not stand in for
model-discovery evidence. The gap this rule exists to close, found live: asked in plain words to
"add the skill I saved in your folder," Haiku invented a nonexistent marker syntax
({"@ag.embed": {"@ag.references": ...}}) and the engine accepted it as literal data — a failure
none of the coached cells (including matrix_w7.py, whose prompt names @ag.file outright)
could ever have caught, because they never test whether the model reaches for the real mechanism
on its own. resources/matrix_g1_guidance_discovery.py is this directory's first mechanism-blind
cell, promoted after the platform-guidance fix closed that exact gap; it reuses qa_matrix_lib.py
the same way the separate one-shot benchmark (Tier B) does — check there before writing a new
mechanism-blind cell from scratch, to avoid duplicating scaffolding.
Session control cells
resources/session_control.py is a second, standalone driver: sixteen cells that cover Stop,
durable commands, and the runner's recovery paths (owner release, park/resume, watchdog
quarantine). It drives the same product endpoint and asserts on the same wire, but it needs its
own account bootstrap, so it runs as a separate process rather than as qa_product.py cells. See
resources/path_triggers.py for the exact mandatory-cell mechanism.
These cells are MANDATORY — run them, not just the standing gate — whenever the release diff
touches any of:
services/runner/src/sessions/**
services/runner/src/engines/sandbox_agent/**
api/oss/src/core/sessions/**
api/oss/src/tasks/asyncio/sessions/**
api/oss/src/apis/fastapi/sessions/**
Run every cell with one line:
uv run resources/session_control.py --cells all --harness pi_core --sandbox local
Add --project <docker-compose project name> to run the eight cells that need direct Docker and
Postgres access (sandbox-gone, records-outage, restart-after-stop, runner-gone,
runner-gone-late, post-stop-row, codex-child, stale-tail) and the abort-log subcheck inside
stop-after-finish. Without --project, those eight cells SKIP with a named reason. The
stop-after-finish HTTP check still runs, but only its abort-log subcheck is unavailable. The
other eight cells
(stop-warm, double-send, stale-stop, stop-approval, stop-after-finish,
repeat-stop, concurrent-stops, stop-during-completion) run over HTTP alone against any
deployment. Add
--resume <path to a prior run's results.json> to pick a lost run back up: any cell already
recorded there is loaded instead of re-run.
Results land in a timestamped folder under ~/agenta-qa-evidence/ (override with
AGENTA_QA_RUNS_DIR), as results.json and summary.md — the same PASS/FAIL/SKIP shape as the
rest of the gate. When a release path makes session control mandatory, pass that artifact to the
standing gate with --session-control-results <path>: a missing or incomplete artifact stops the
gate before the matrix runs, and a recorded FAIL makes the final gate exit nonzero.
Environment, by name. Same three-variable discipline as the rest of the gate, no env-file
fallback:
AGENTA_BASE — the deployment origin.
AGENTA_ADMIN_KEY — mints the ephemeral account this driver runs under. Lives in
~/.agenta-qa-secrets.env.
QA_OPENAI_API_KEY — stocked into that account's vault so the pi_core and codex harnesses
have a provider key. Lives in ~/.agenta-qa-openai.env.
ANTHROPIC_API_KEY — only required for --harness claude, stocked into the same vault the
same way. Lives in ~/.agenta-qa-secrets.env. A pi_core- or codex-only run does not need it.
A Daytona run additionally needs a Secrets-capable Daytona key on the runner; the key in most
session env files returns 403 on the Secrets endpoint, so check that before trusting a Daytona
result.
When results lie
The runtime fails open: a component can break, get logged, and the turn still succeeds with a
normal-looking answer. A green turn is therefore not proof on its own. Before trusting a pass,
read resources/LESSONS.md — every trap there produced a green test that proved nothing. The three
that bite hardest: replay conversation history byte-faithfully (tool parts included) or every turn
silently goes cold; re-run any prior blocker-level finding after a redeploy before believing it;
and a multi-turn check that never leaves the warm daemon, on a deployment with no object store,
proves nothing about the durable working directory (LESSONS #16).
Resources (read on demand)
resources/coverage.md — the cells (harness × sandbox × auth), the journeys (chat, mount, tool,
approve, deny, commit, warm, cold1, cold2, mcp) with a one-line meaning for each, the continuity
tiers and their method, and a table of what each cell needs beyond the three env vars.
resources/LESSONS.md — the traps. Read before writing or trusting any agent QA test.
resources/qa_product.py — the gate driver (cells × journeys).
resources/matrix_gw1_gateway_tools.py — [coached, with one mechanism-blind leg] the
gateway tool surface against a real provider: search filters by policy (the denied key never
reaches the model), an allowed tool executes unattended with a genuine provider result, and an
ask-tier tool parks with the right stored identity and is answered through the interactions
API — the durable plane a reloaded browser uses, which no other cell covers. The fixed
matrix proves approvals with a builtin, so nothing else notices when the compiled policy and
the enforced policy drift apart. Defaults to the no-auth text_to_pdf connection; --integration
and --connection point it elsewhere. SKIPs, naming the fixture, when no valid connection
exists, and SKIPs rather than failing the release when the model provider itself errors.
Made mandatory by the gateway rule in path_triggers.py.
resources/path_triggers.py — the path-scoped rules: one dict mapping a path glob to the cells
a release must run when its diff touches that glob, plus the two functions the driver calls.
Runs standalone as a preview (--release-base <ref>). Add a rule here whenever new coverage is
only meaningful for changes in one part of the tree.
resources/qa_probe.py — a one-turn wire probe: uv run resources/qa_probe.py confirms the
product path answers at all before running the full gate.
resources/qa_commit_approval.py — [coached] the mandatory pre-handoff commit-approval
round trip (see above). Self-contained; does not import qa_product.py.
resources/qa_matrix_lib.py — shared helpers (session/turn plumbing, workflow/revision REST
calls, the multi-round approval loop) for the matrix_w*.py adversarial cells below. Import
only, no CLI. It also holds the two cross-cutting invariants every cell should fold into its
verdict — check_no_blank_success_on_refusal and check_no_silent_turn (see below). If you
write a new cell, wire check_no_silent_turn into its PASS condition.
resources/matrix_w3.py — [coached, with a narrow mechanism-blind sliver] two sessions,
disjoint edits; session B is given a stale base_revision_id and ZERO coaching on recovery.
Two-tier pass: autonomous correct recovery passes outright; a model that diagnoses the 409
correctly and asks before re-attempting a config write ALSO passes (that is desirable caution,
not a failure) once one bare "yes, retry" permission (no mechanics) completes the recovery.
Guards both the optimistic-concurrency base check and the instructive-error design (a 409 must
be readable without coaching). The initial action in both sessions is still coached — only the
RECOVERY step is mechanism-blind; don't cite this cell for "the model finds commit_revision
unprompted."
resources/matrix_w4.py — [coached] a pending approval whose base goes stale while it
waits (a second session commits first); the EXECUTE-time check must catch it, not just the
gate-time check.
resources/matrix_w5.py — [coached] interrupt a running turn (steer) then use the session
again. Caught a real bug: the durable mount never gets re-established after a steer, breaking
every subsequent turn on that session. Distinct from Mahmoud's own steer repro (that one is a
turn-currency/heartbeat bug; this is a mount-lifecycle bug) — keep the two separate when
triaging.
resources/matrix_w7.py — [coached] an agent writes a workspace file and commits it via an
@ag.file marker; asserts the approval manifest carries digest+bytes and the commit lands with
the exact bytes. Caught a real bug (fixed, PR #5763-adjacent runner fix): the gate approved
cleanly but execution always refused with authorization_missing — no file-marker commit could
land. Re-verified PASS after the fix. Note: DEFERRED_NOT_EXECUTED on a queued second tool
call is benign (another gate is already pending), not a real tool error — don't let it fail
this cell. Naming @ag.file verbatim in the prompt is correct for this cell's backend-path
purpose but means it cannot catch a model failing to find the marker syntax unprompted — see
Tiers above.
resources/matrix_w1_daytona.py — [coached] the commit round-trip on sandbox=daytona
(every other cell runs local). Needs a funded provider vault key (see the file's own docstring
for the exact vault-secret shape and the provider_key slug gotcha). PASS confirms
DaytonaWorkspaceReader and the placeholder-secrets flow live.
resources/matrix_t8_saved_files.py — [coached] T8: the Daytona remote agent mount (the
durable agent-files/ folder the playground file drawer writes into), which needs the ngrok
tunnel the runner discovers at sandbox-acquire time. Writes a marker file into the mount via
the mounts API (POST /mounts/agents/sign + PUT /mounts/{id}/files) before the run, then
asserts the runner log line remote agent mount active for artifact=<id> appears, a tool
output actually carries the real marker content (not hallucinated), and the commit lands with
it. Verified PASS 3/3 runs after the tunnel-seat fix (2026-08-06); this line was absent on
every attempt before that fix landed.
resources/matrix_t9_agent_tools.py — [coached] T9: the runner restores the agent's own
tools from agent-files/.tools/ before a session (agent-tools-setup.ts). Plants a
setup.sh and a bin/qa-tool through the mounts API, opens a fresh session, and asserts the
first tool call is the exact probe and its output payload carries both planted tokens, plus a
line setup.sh appended to agent-files/.tools/runs.log, read back through the mounts API
with no model in the loop, plus the agent_tools_setup stage for THIS session in the runner
log when --runner-container is given. --sandbox local|daytona, --harness pi_core|claude.
Mandatory (via path_triggers.py) when the restore step or the sandbox image recipes change.
resources/matrix_w7_per_harness.py — [coached] matrix_w7.py's exact scenario run
identically on all three harnesses (claude, codex, pi_core), each classified PASS/FAIL/SKIP
independently. Exists because W7 originally ran on Claude only, and that scenario-coverage gap
is exactly what let the Codex approve-then-fail P0 (2026-08-06) ship: commit_revision + file
marker + HITL approval on a harness this suite never exercised. claude uses subscription auth
(no vault dependency); codex and pi_core need a funded OpenAI provider_key vault secret
(mirrors cells X1/C3 in qa_product.py) and correctly SKIP with the exact reason when it's
missing or ambiguous — a SKIP here is an untested harness, not a pass, and must be named as such
in any release summary. --only <harness> runs a single leg without re-spending budget on the
others. Verified PASS on claude (2026-08-06). codex/pi initially SKIPPED on this shared preview
stack because its vault held zero — then, transiently, an ambiguous multiple — OpenAI candidates
(other concurrent agents' activity on the same shared project); resolved 2026-08-06 by stocking
one unambiguous OpenAI provider_key secret (no stale entries existed to remove). Re-verified
PASS on codex after stocking the key (session 58ce3a58-8d04-40ac-99e4-c44eaa5d7b06).
resources/matrix_w7_daytona.py — [coached] matrix_w7.py's exact scenario with
sandbox=daytona instead of local. The local-only original W7 is exactly why this bug hid: the
Daytona transport rejects NUL bytes in argv, which the @ag.file manifest walk was emitting, so
no workspace-file commit could EVER land on Daytona, on any harness, until the 2026-08-06 fix
(found during the same P0 triage, live-verified twice on codex+Daytona: sessions f3fa4335,
f2f22056). This cell is the sandbox-axis regression guard, staying on the claude harness. Needs
the same funded Anthropic vault key as matrix_w1_daytona.py. Verified PASS (2026-08-06).
resources/matrix_invariant_commit_auth_refusal.py — [coached scenario; the invariant itself
is mechanism-level] the generic invariant: no tool_result with empty output and
isError:false may exist for a call whose runner log says [commit-auth] refused (the
silent-blank-success class — the P0's actual failure shape, distinct from the scenario-coverage
gap matrix_w7_per_harness.py addresses). The check itself
(qa_matrix_lib.check_no_blank_success_on_refusal) reads the runner's own log line against the
wire outcome and is meant to be reusable by any cell that exercises marker-carrying commits, not
just this one. This cell's trigger is best-effort: run W7's flow, let the legitimate commit
consume its authorization record, then REPLAY the byte-identical approval-carrying request
(a duplicate submission) hoping to force a second, doomed authorizeExecution attempt. Verified
2026-08-06: the replay did NOT reproduce a refusal (the runner treats the identical replayed
history as already-resolved and answers conversationally instead of re-invoking the tool), so
the cell correctly SKIPped rather than claiming a false pass on an invariant it never exercised.
A reliable deterministic trigger (e.g. a genuine cold-resume stale-approval replay, or a crafted
duplicate toolCallId) is an open follow-up; until then, treat any SKIP from this cell as "the
invariant was not tested this run," never as green.
qa_matrix_lib.check_no_silent_turn(turns) — [mechanism-level invariant; no cell of its
own] no turn may come back completely bare: no text, no tool call, no approval gate, no file
or data payload, and no error. That combination is a swallowed provider failure (ASD-EST100) —
the model call is rejected, the error is dropped on the way back, and the turn is reported as a
clean empty finish, so the user sees a blank bubble with no reason anywhere. It matters most in
cells whose PASS depends on something NOT appearing (no error, no leak, no blank success): a
turn that produced nothing satisfies those by doing nothing at all. Its content definition
deliberately mirrors content_parts_emitted in the product's own Vercel egress
(sdks/python/agenta/sdk/agents/adapters/vercel/stream.py) — reasoning does NOT count, so a
turn that only thought is still a violation. Wired into matrix_w7*.py, matrix_t8_saved_files.py,
matrix_b1_builtin_find.py, matrix_invariant_commit_auth_refusal.py,
matrix_l3_abandoned_approval.py, matrix_w3.py, matrix_w4.py and matrix_w5.py as
... and not silent["violations"]; add the same conjunct to any new cell. The one thing a cell
must exclude itself is a turn it deliberately aborted or interrupted, which legitimately ends
bare — matrix_w5.py shows the pattern (it checks only its post-interrupt turns).
resources/matrix_g1_guidance_discovery.py — [mechanism-blind] does the platform guidance
actually change what the model does? The trial prompt is Mahmoud's own verbatim phrasing from
the live session that found the bug ("can you add gstack-autoplan skill to your skills (i saved
it in your folder)") — no tool, operation, or marker syntax named. Before the guidance existed,
this exact phrasing made a model copy the skill into its own harness-local skills folder and
claim success without ever proposing a commit, 3/3 live (session b59cb549). Two parts per
harness/model/sandbox leg: PROBE reads the rendered instructions file out of the workspace and
asserts the fenced platform-guidance block and the skill-location sentence are really there;
TRIALS run the prompt N times, PASS only when a commit_revision gate fires, is approved, and the
STORED revision carries the skill (copying into the harness's own folder is a FAIL). Setup
gotchas baked into qa_matrix_lib.py (PI_CORE_HARNESS_KIND, PI_CORE_HAIKU_MODEL): the Pi
harness kind enum is "pi_core" (bare "pi" 500s), and pi_core rejects a bare "haiku" model
id (needs "claude-haiku-4-5"); codex accepts its curated short alias ("gpt-5.6-luna") bare.
Both legs need sandbox=daytona + a vault key (codex: OpenAI; pi_core: the same Anthropic key
matrix_w1_daytona.py documents) and SKIP with the exact reason when the credential is missing
or ambiguous. --only <leg> runs a single leg without re-spending trial budget on the other.
Neither leg has yet scored a clean 4/4 live, and both failure modes are real, reproducible model
mechanics misses, not noise or infra flake — treat this cell's discovery rate as a genuine open
quality question, not a settled pass:
- pi_core: 2/3 (2026-08-06). Trial 2's model ran
cp -r agent-files/gstack-autoplan .agenta-imports/ (copying the whole directory) then referenced a marker path that didn't
match where the file landed; the engine correctly denied it fail-closed
(approved-content resolution failed ...: gstack-autoplan/SKILL.md does not exist under .agenta-imports/.; deny). Not a product bug — the deny is doing its job — but evidence the
model fumbles the exact copy-then-reference mechanics some of the time.
- codex: 2/3 (2026-08-06, re-verified after stocking a clean OpenAI vault key — the earlier
SKIP was purely the missing/ambiguous credential, now resolved). Trial 1 didn't attempt the
mechanism at all: zero gates approved, the model just replied "I can't add skills by simply
placing files in the skills folder; skills must be enabled through the agent configuration"
(session 4fa17164-3ada-4ef6-86b5-63bf1c64f10b) — it named the right concept but never called
read_config/commit_revision to act on it. Trials 2 and 3 passed cleanly (sessions
4c6a758b-39a3-4789-890e-06f3d68196b6, bf8a1283-667c-45c7-ab0d-b112e12106db).
Worth a call: whether the guidance text needs to be more directive (e.g. explicitly say "copy
the FILE, not the directory" and "always attempt the tool calls, don't just describe the
mechanism"), or whether ~2/3 is an acceptable bar for this feature's launch.
Builtin capability cells (matrix_b*.py) — does the harness's own tooling still work
resources/matrix_b1_builtin_find.py — [coached, harness-mechanism test] one native
file-search call per harness (write three known marker files, then ask the model to locate
them via its own search capability — never a manual directory listing), asserting the exact
filenames come back in the TOOL OUTPUT payload, never the reply. Exists because nothing else in
the gate ever exercised a harness builtin: verify-runner's overnight diagnosis found Pi's
find builtin dead 52/52 across two benchmark runs (it shells out to the vendored fd binary
with a flag that only exists from fd 9 onward; the runner image ships fd 8.6.0) — a total
capability loss that sat invisible with nothing calling it. This cell closes that discoverability
gap for the class, not just this one instance. Open discrepancy, not yet reconciled: this
cell's own pi_core leg PASSED twice, live, with real filenames back
(2026-08-07, sessions dd9c51ef-92d0-4780-855e-da6f48e07d9f and
47f02ab1-61ec-4db8-a0d6-2d96e98b9188) — which does not match "52/52 failed". Either the break is
conditional on a flag/option this cell's simple case never exercises, or something already
changed; needs reconciling with verify-runner before Pi's find gets called either fixed or
still broken. Codex SKIPs by design, not tested: its exec output doesn't land in the
tool-output-available payload's .output field (the same quirk qa_product.py's j2_mount
already names and skips codex for), so this cell's evidence extraction cannot see codex's real
results — a codex-shaped extraction is a follow-up. Claude PASSED cleanly on the two-turn
version (session 4081e9ee-11c1-4f12-8245-6c390f90e9d8). Needed two EXPLICIT turns (write, then
search) — one combined instruction left claude stopping after the write step without
attempting the search at all.
The lifecycle cells (matrix_l*.py) — cold ↔ warm, and what survives each transition
These four cover the session-lifecycle work: which config changes are applied to a RUNNING
sandbox and which tear it down, and what happens to a pending approval, a client tool and the
durable mount across each transition. They all assert the STORED turn ledger
(POST /sessions/turns/query → one sandbox_id per turn) or the STORED interaction rows
(POST /sessions/interactions/query), never the SSE echo — nothing about warm-versus-rebuilt
ever reaches the stream. An empty ledger FAILS a cell; missing evidence is not evidence.
resources/matrix_l1_lifecycle_routes.py — MANDATORY. [mechanism-blind] the routing matrix
itself: for each kind of mid-conversation config change, assert the route the runner took. One
sandbox id = applied in place, two = rebuilt. Blocks on six cases: no change must stay warm; an
instructions edit, a permissions edit and a tool-catalog edit must escalate; and a
same-connection model switch must stay warm on BOTH claude and pi_core. The pi_core model case
(added 2026-08-29) is the standing trap for the wire-spelling bug class: the router once keyed
its table on the bare "pi" literal while the wire carries "pi_core", every playground model
switch silently rebuilt, and the claude-only case could not see it (#6364). This is the cell
that would have caught the cold1 rot described below.
resources/matrix_l2_approval_across_config_change.py — MANDATORY. [coached] the killer
combination: an approval answered while a config change rides along in the SAME request. It is
the regression test for the applied-state bug (the pool used to stamp the INCOMING fingerprint
on the approval-resume path, so the next turn continued warm on an environment running
something else). Asserts the gated commit lands, the approval row ends resolved/responded,
and — the real tell — the config change is not swallowed: it takes effect on the FOLLOWING
turn, via a rebuild, in both the instructions and the permissions variant.
resources/matrix_l3_abandoned_approval.py — MANDATORY. [coached] the user sends a new
message instead of answering the card. Asserts the gated tool does NOT run (an unanswered
approval is not consent), the row is swept to cancelled rather than left pending, and the
session still works. cancelled vs pending is the loud-vs-silent distinction: a pending row
is a card sitting on the page that no process is waiting on.
resources/matrix_l5_live_route_observed.py — MANDATORY. [mechanism-blind, with a control]
the other half of L1: an instructions edit made mid-conversation must actually be OBSERVED by
the harness, not merely written to disk. Runs the same configuration on a fresh cold session as
a control, so a failure isolates the runner rather than blaming the model; when the control also
fails it reports INCONCLUSIVE instead of a confident wrong verdict. It asserts the edit, never
the route, so it stays meaningful if the facet is ever made live again. Failed 2026-08-06
(claude/local) and now passes — see the finding note below. Extended 2026-08-06 (overnight
gate run) to all three harnesses in one invocation (--only <harness> for a single leg) —
this MANDATORY blocker cell had only ever run on claude; codex and pi_core were a named gap.
Verified PASS on all three the same night the harness matrix landed.
resources/matrix_l4_client_tool_lifecycle.py — nice-to-have. [coached] the client-tool
round trip, and the only cell that covers client tools at all. Asserts the browser's result
reaches the model and the client_tool interaction is stored. It RECORDS rather than asserts
the sandbox count, which is two today: a client-tool pause is deliberately not parkable
("warm-hold": RESERVED, not built, #5384), so every client-tool round trip currently costs a
rebuild. If that number ever reads one, the warm hold landed and the docstring needs updating.
resources/matrix_n1_session_context.py — MANDATORY. [journey, with two controls] the
per-turn session facts on the path the PRODUCT uses. Renames a session between two turns and
asks the agent for the name, asks the agent for its own display name, and posts a forged
meta.session_context that the service must ignore. Every expected value carries a random
token minted for the run and spoken nowhere in the conversation, so a transcript-derived
answer cannot match; the second ask is the exact shape of #6661, because by then the FIRST
name is in the transcript and the current one is only in the stored header. This is the one
cell that asserts on model prose, and it does so because nothing else can: turnContext is a
prompt string on the service-to-runner payload, the runner logs nothing, and it is
deliberately kept out of request.messages and out of persisted input. Two controls keep a
FAIL honest — an echo probe (a model that cannot repeat a literal token makes the cell report
INCONCLUSIVE) and a read-back of the stored header after every rename. Mandatory (via
path_triggers.py) when the SDK session-context module, the platform-prompt renderer, the
agent handler, or the API's session-context resolver changes. The gate could not see this
class at all before v0.115.3: #6661 shipped green because the API stamped the facts in its
invoke prelude, which never runs for a playground turn, and no cell renamed a session
mid-conversation or asserted on the facts.
Finding the lifecycle cells surfaced (2026-08-06, claude on local, reproduced 3×) — FIXED: the
workspaceFiles live route rewrote the instruction file and advanced applied state, but the
running harness never re-read it. A warm session kept obeying the instructions it started with
while the pool reported the NEW fingerprint, so every later turn matched and continued warm and the
user's edit had no effect until something else evicted the session. A cold session with the
identical configuration obeyed it immediately, which is what isolated the runner. This was the
failure desired-state.ts refuses to allow for the prompts facet ("refreshing them and claiming
the model saw the change would be a lie") reappearing on the facet that WAS made live — and note
the direction: before the live route existed, an instructions edit forced a rebuild and therefore
took effect on the next turn, so it was a regression in what the user sees, not a speedup.
matrix_l5_live_route_observed.py is the repro.
The fix withdrew the route: workspaceFiles now routes to rebuild-sandbox in the capability
table, and refresh-workspace left LIVE_ACTION_KINDS so restoring the table alone fails closed.
An instructions edit costs a sandbox again, which is what it cost before the optimisation. L1's
instructions case therefore expects TWO sandbox ids, and L2's first variant expects two as
well — if you are reading an old green from before 2026-08-06, those cells expected one. The
intended next shape is refresh THEN reopen the session, which needs the reopen to build its
session init from the incoming request first, and needs proving on L5 rather than asserting.
Why cold1 changed (read before trusting an old green): that tier used to force its eviction
by editing instructions.agents_md, which the lifecycle work briefly made a LIVE route — the tier
would have gone on passing while measuring warm reuse, and park's "one sandbox id is meaningful"
argument rests on cold1 reporting two on the same deployment. It now moves harness.permissions
(the harnessSession facet → reopen-session, deliberately not live) and ASSERTS two distinct
sandbox ids. It sta
…(truncated)
1---2name: agent-release-gate3description: Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an agent-workflows release, or after changing the runner, the SDK agent adapters, the runner Docker images, or the agent service. Triggers: "run the release gate", "QA the agent runtime", "does the agent still work end to end", "pre-release agent QA".4---56# Agent release gate78Product-level sanity QA for the agent runtime, one layer below the playground UI. The question9is not "is every detail right" — it is "**if a user opens the product and does the obvious first10things, do they work?**" This is the gate a release passes before shipping.1112Every check asserts on the **wire** (the SSE frame types the browser sees) and on **side effects**13(the file really persisted, the revision really incremented) — never on what the model says. That14makes it deployment-agnostic: point it at any stack and the assertions still hold.1516## Run it1718Set three environment variables for the deployment under test, then run the gate:1920```bash21export AGENTA_BASE=https://your-stack.example.com # deployment origin22export AGENTA_PROJECT_ID=... # target project23export AGENTA_API_KEY=... # project API key2425uv run resources/qa_product.py --all --custom-slug <vault-slug> --custom-name "<display-name>" --require-store # everything26uv run resources/qa_product.py --cell P1 # one cell27uv run resources/qa_product.py --cell C1 --only chat # one journey28uv run resources/qa_product.py --cell S2 --only warm --only cold1 --require-store # continuity29```3031**EXPORT the three variables, do not just set them.** The driver falls back to an env FILE when32`AGENTA_*` is absent from its environment, which is helpful interactively and dangerous in a33release run: a credentials file of bare `KEY=value` lines sourced with `. file` sets the shell34only, the child `uv run` process inherits nothing, and the driver silently runs the whole gate35against WHATEVER DEPLOYMENT the fallback file names. The failure surfaces as `401 Invalid36credentials` from a stage whose key you just watched answer 200, or worse as a green run37recorded against the wrong stack. Use `set -a` around the source, or `export` each variable,38and confirm the stage in the results before trusting them. (Cost a staging gate run on392026-08-28; the fallback degrades to "wrong deployment", never to "no credentials".)4041Paths are relative to this skill's directory. The deployment's vault must hold the provider keys42the cells use (Anthropic / OpenAI / OpenRouter). If the three env vars are unset the driver stops43immediately and names exactly what is missing; a legacy `--env-file <path>` fallback also exists.44`--all` includes cells P2, P2b, and P3 (a custom OpenAI-compatible provider, with P2 and P2b45running locally and P3 on Daytona), which need a vault slug passed via `--custom-slug`; the driver fails46fast if it's missing. `--custom-name` is also required because `model_keys` is built from the47display name rather than the stable slug. Custom-provider cells skip the credential-rotation48journey because their value is write-only and cannot be safely restored. Cells S1, S2 and C149additionally need50the subscription sidecar logged in on the target deployment — see `resources/coverage.md` for what51each cell requires. The Daytona cells (C2, C4, P3, X2) additionally run the `secret_opaque`52journey and need the runner's Daytona API key to manage Secrets, because credential hiding is on53by default; without it those cells fail at sandbox creation with an error naming the permission.5455**The one flag a release conductor must not skip past.** The continuity journeys56(`warm`, `cold1`, `cold2`) only mean anything on a **store-backed** deployment: with no object57store the runner degrades silently to an ephemeral working directory, so those journeys SKIP by58default and FAIL with `--require-store`. Run the gate against a deployment with `AGENTA_STORE_*`59configured and pass `--require-store`, or the greenest possible run still says nothing about60durability. `cold2` additionally needs an operator hook that SIGKILLs the runner replica61(`--cold2-replace-cmd`) and SKIPs without it.6263**The flag that makes the gate fit the release: `--release-base`.** The matrix is fixed, so64without it a release that reworked a subsystem gets exactly the coverage of a release that did65not touch it. Pass the ref the release branches from, and the driver reads the release's own66changed paths, matches them against the rules in `resources/path_triggers.py`, and makes the67cells those rules name MANDATORY for this run:6869```bash70uv run resources/qa_product.py --all --release-base origin/main --require-store # every release run71uv run resources/path_triggers.py --release-base origin/main # preview only, runs nothing72```7374A mandatory cell that lives in `qa_product.py` is added to the run even when `--cell` did not ask75for it. A mandatory cell that is a standalone `matrix_*.py` script is a separate process the76driver cannot observe, so it is printed, written to `mandatory.json`, and listed in `summary.md`77under "Mandatory for this release" — **the release is not green until each of those has a recorded78result of its own.** If a rule names a cell that does not exist, the driver stops before running79anything and says so: the release changed code a rule protects and the coverage was never80written, which is the one outcome that must never read as green. Add `--changed-path` to state81paths by hand where the checkout is not the release branch. With no `--release-base` and no82`--changed-path` nothing changes, so every existing invocation behaves exactly as before.8384Adding a rule is one line in `PATH_TRIGGERS` (a glob, and the cells it makes mandatory) plus the85cell it names. Matching is `fnmatch` over the whole repo-relative path, so `*` crosses directory86separators and `a/b/*` covers the whole subtree; write `**` so a subtree rule reads as one.8788**Reading the result.** Each journey prints `PASS`, `FAIL`, or `SKIP` with a one-line reason, and89a per-cell markdown table lands with the full JSON in `./qa-gate-runs/<timestamp>/` (override the90location with `AGENTA_QA_RUNS_DIR`). Runs are written to the current working directory, never into91the skill. `SKIP` is expected where a journey does not apply to a cell (for example `mcp` on any Pi92cell — user MCP is Claude-only). Any `FAIL` blocks the release until triaged.9394**A SKIPped integration test in a security- or concurrency-bearing area is a FAILURE in your95summary line, not green.** State it as "N passed, M skipped OF WHICH k are untested claims" and96name the k. A commit-lock race test skipping for want of a reachable Postgres is exactly how a97one-line syntax error (`SET LOCAL lock_timeout` with a bind parameter, which Postgres rejects98outright) survived 1911 green tests before a human hit it as his first live action.99100**The two journeys that run many things at once: `burst` and `crosstalk`.** Every other journey101drives one run at a time, so the gate only ever saw faults that reproduce on a quiet deployment.102The credential-delivery fault of AGE-4249 does not: about one production first message in five103failed because some fresh Daytona sandboxes start without their Secret substitution wiring, and104per cold sandbox that is roughly an 8 percent fault. `burst` sends 16 first messages at the same105time on 16 brand new sessions, so the run buys 16 cold starts instead of one. `crosstalk` runs 3106two-turn conversations with long output beside 2 approval flows, and checks that no stream carries107another session's nonce, except on the codex harness, where the gate rides a platform tool with108empty arguments, so the approval command carries no nonce and isolation is not checked there109(`nonce_checked=false`). Both are Daytona-only by default and skip elsewhere; both report the110runner's stable error code per run, so a `credential_delivery_failed` names itself.111112```bash113uv run resources/qa_product.py --cell C4 --only burst --only crosstalk114uv run resources/qa_product.py --cell C4 --only burst --burst-size 24 # more cold starts115uv run resources/qa_product.py --cell C3 --only crosstalk --concurrency-everywhere # local too116uv run resources/test_qa_product_concurrency.py # offline tests117```118119**Read a green burst honestly.** At an 8 percent per-cold-start fault rate, 8 runs miss the fault12051 percent of the time, 16 miss it 26 percent, and two Daytona cells at 16 miss it about 7 percent.121A PASS is a sample, not an all-clear, and the result says so in its own reason line. A FAIL is122proof.123124Each concurrent run holds its own Daytona sandbox, about 5 GiB of the organization's disk, and a125parked sandbox keeps counting until its auto-delete window closes, so a burst of 16 is about 80126GiB in flight. The counts are `--burst-size` (default 16), `--crosstalk-conversations` (default 3)127and `--crosstalk-approvals` (default 2). The cap is 32 concurrent runs: 32 for the burst size, and12832 for the two crosstalk counts TOGETHER, because what costs disk is what runs at once. When the129provider refuses on capacity the journey reports SKIP with a loud reason, never a PASS or a FAIL,130because nothing about the product was measured. `--concurrency-timeout` (default 300s) bounds one131TURN and rides into the stream as an absolute deadline, so a two-turn crosstalk run gets twice132that and a stream that never ends is abandoned rather than followed.133134A release that changes `services/runner/src/engines/sandbox_agent/**` or135`services/runner/src/providers/daytona*` makes the Daytona cells C2, C4 and X2 mandatory through136`path_triggers.py`, and forces `burst` and `crosstalk` into the run even when `--only` named137something else. That is how these journeys reach a release that needs them.138139**Before a human gets a deployment URL, run `resources/qa_commit_approval.py` too.** It is not140part of `qa_product.py`'s cell × journey matrix — none of that matrix's journeys drive a live turn141against a REAL, saved workflow revision (the `commit` journey only exercises the REST API; `chat`,142`tool`, `approve`/`deny` all run against an inline, unsaved config), so nothing else in the gate143observes the S3b single-use execution-authorization gate actually firing around a real config144mutation. This script does: create a real workflow + revision, invoke a live agent turn that calls145`read_config` then `commit_revision`, expect the pause, approve it in-band, and verify the new146revision landed with REST fetch-back. Treat a FAIL here as blocking, the same as any other gate147FAIL.148149## Tiers: coached vs. mechanism-blind150151Every cell in this gate (`qa_product.py`'s journeys, `qa_commit_approval.py`,152`qa_probe.py`, and all `matrix_w*.py` cells) declares which of two tiers it belongs to, in its153own docstring:154155- **Coached (backend-path test).** The prompt names the mechanism verbatim — which tool, which156 operation, which target path. This proves OUR CODE works (the gate, the base check, the157 authorization handoff, the sandbox path) when the right call is made. It proves NOTHING about158 whether a model finds that call from a plain-language human ask.159- **Mechanism-blind (model-behavior test).** The prompt is phrased the way a real user types —160 no tool names, no operation names, no schema hints. Only these cells license a claim about161 what the model can do unprompted.162163**Rule: claims of model behavior may only cite mechanism-blind cells.** Most cells in this164directory are coached tier — they test backend paths, correctly, but do not stand in for165model-discovery evidence. The gap this rule exists to close, found live: asked in plain words to166"add the skill I saved in your folder," Haiku invented a nonexistent marker syntax167(`{"@ag.embed": {"@ag.references": ...}}`) and the engine accepted it as literal data — a failure168none of the coached cells (including `matrix_w7.py`, whose prompt names `@ag.file` outright)169could ever have caught, because they never test whether the model reaches for the real mechanism170on its own. `resources/matrix_g1_guidance_discovery.py` is this directory's first mechanism-blind171cell, promoted after the platform-guidance fix closed that exact gap; it reuses `qa_matrix_lib.py`172the same way the separate one-shot benchmark (Tier B) does — check there before writing a new173mechanism-blind cell from scratch, to avoid duplicating scaffolding.174175## Session control cells176177`resources/session_control.py` is a second, standalone driver: sixteen cells that cover Stop,178durable commands, and the runner's recovery paths (owner release, park/resume, watchdog179quarantine). It drives the same product endpoint and asserts on the same wire, but it needs its180own account bootstrap, so it runs as a separate process rather than as `qa_product.py` cells. See181`resources/path_triggers.py` for the exact mandatory-cell mechanism.182183**These cells are MANDATORY** — run them, not just the standing gate — whenever the release diff184touches any of:185186- `services/runner/src/sessions/**`187- `services/runner/src/engines/sandbox_agent/**`188- `api/oss/src/core/sessions/**`189- `api/oss/src/tasks/asyncio/sessions/**`190- `api/oss/src/apis/fastapi/sessions/**`191192Run every cell with one line:193194```bash195uv run resources/session_control.py --cells all --harness pi_core --sandbox local196```197198Add `--project <docker-compose project name>` to run the eight cells that need direct Docker and199Postgres access (`sandbox-gone`, `records-outage`, `restart-after-stop`, `runner-gone`,200`runner-gone-late`, `post-stop-row`, `codex-child`, `stale-tail`) and the abort-log subcheck inside201`stop-after-finish`. Without `--project`, those eight cells SKIP with a named reason. The202`stop-after-finish` HTTP check still runs, but only its abort-log subcheck is unavailable. The203other eight cells204(`stop-warm`, `double-send`, `stale-stop`, `stop-approval`, `stop-after-finish`,205`repeat-stop`, `concurrent-stops`, `stop-during-completion`) run over HTTP alone against any206deployment. Add207`--resume <path to a prior run's results.json>` to pick a lost run back up: any cell already208recorded there is loaded instead of re-run.209210Results land in a timestamped folder under `~/agenta-qa-evidence/` (override with211`AGENTA_QA_RUNS_DIR`), as `results.json` and `summary.md` — the same PASS/FAIL/SKIP shape as the212rest of the gate. When a release path makes session control mandatory, pass that artifact to the213standing gate with `--session-control-results <path>`: a missing or incomplete artifact stops the214gate before the matrix runs, and a recorded FAIL makes the final gate exit nonzero.215216**Environment, by name.** Same three-variable discipline as the rest of the gate, no env-file217fallback:218219- `AGENTA_BASE` — the deployment origin.220- `AGENTA_ADMIN_KEY` — mints the ephemeral account this driver runs under. Lives in221 `~/.agenta-qa-secrets.env`.222- `QA_OPENAI_API_KEY` — stocked into that account's vault so the `pi_core` and `codex` harnesses223 have a provider key. Lives in `~/.agenta-qa-openai.env`.224- `ANTHROPIC_API_KEY` — only required for `--harness claude`, stocked into the same vault the225 same way. Lives in `~/.agenta-qa-secrets.env`. A pi_core- or codex-only run does not need it.226227A Daytona run additionally needs a Secrets-capable Daytona key on the runner; the key in most228session env files returns 403 on the Secrets endpoint, so check that before trusting a Daytona229result.230231## When results lie232233The runtime **fails open**: a component can break, get logged, and the turn still succeeds with a234normal-looking answer. A green turn is therefore not proof on its own. Before trusting a pass,235read `resources/LESSONS.md` — every trap there produced a green test that proved nothing. The three236that bite hardest: replay conversation history byte-faithfully (tool parts included) or every turn237silently goes cold; re-run any prior blocker-level finding after a redeploy before believing it;238and a multi-turn check that never leaves the warm daemon, on a deployment with no object store,239proves nothing about the durable working directory (LESSONS #16).240241## Resources (read on demand)242243- `resources/coverage.md` — the cells (harness × sandbox × auth), the journeys (chat, mount, tool,244 approve, deny, commit, warm, cold1, cold2, mcp) with a one-line meaning for each, the continuity245 tiers and their method, and a table of what each cell needs beyond the three env vars.246- `resources/LESSONS.md` — the traps. Read before writing or trusting any agent QA test.247- `resources/qa_product.py` — the gate driver (cells × journeys).248- `resources/matrix_gw1_gateway_tools.py` — **[coached, with one mechanism-blind leg]** the249 gateway tool surface against a real provider: search filters by policy (the denied key never250 reaches the model), an allowed tool executes unattended with a genuine provider result, and an251 ask-tier tool parks with the right stored identity and is answered through the **interactions252 API** — the durable plane a reloaded browser uses, which no other cell covers. The fixed253 matrix proves approvals with a builtin, so nothing else notices when the compiled policy and254 the enforced policy drift apart. Defaults to the no-auth `text_to_pdf` connection; `--integration`255 and `--connection` point it elsewhere. SKIPs, naming the fixture, when no valid connection256 exists, and SKIPs rather than failing the release when the model provider itself errors.257 Made mandatory by the gateway rule in `path_triggers.py`.258- `resources/path_triggers.py` — the path-scoped rules: one dict mapping a path glob to the cells259 a release must run when its diff touches that glob, plus the two functions the driver calls.260 Runs standalone as a preview (`--release-base <ref>`). Add a rule here whenever new coverage is261 only meaningful for changes in one part of the tree.262- `resources/qa_probe.py` — a one-turn wire probe: `uv run resources/qa_probe.py` confirms the263 product path answers at all before running the full gate.264- `resources/qa_commit_approval.py` — **[coached]** the mandatory pre-handoff commit-approval265 round trip (see above). Self-contained; does not import `qa_product.py`.266- `resources/qa_matrix_lib.py` — shared helpers (session/turn plumbing, workflow/revision REST267 calls, the multi-round approval loop) for the `matrix_w*.py` adversarial cells below. Import268 only, no CLI. It also holds the two cross-cutting invariants every cell should fold into its269 verdict — `check_no_blank_success_on_refusal` and `check_no_silent_turn` (see below). **If you270 write a new cell, wire `check_no_silent_turn` into its PASS condition.**271- `resources/matrix_w3.py` — **[coached, with a narrow mechanism-blind sliver]** two sessions,272 disjoint edits; session B is given a stale `base_revision_id` and ZERO coaching on recovery.273 Two-tier pass: autonomous correct recovery passes outright; a model that diagnoses the 409274 correctly and asks before re-attempting a config write ALSO passes (that is desirable caution,275 not a failure) once one bare "yes, retry" permission (no mechanics) completes the recovery.276 Guards both the optimistic-concurrency base check and the instructive-error design (a 409 must277 be readable without coaching). The initial action in both sessions is still coached — only the278 RECOVERY step is mechanism-blind; don't cite this cell for "the model finds commit_revision279 unprompted."280- `resources/matrix_w4.py` — **[coached]** a pending approval whose base goes stale while it281 waits (a second session commits first); the EXECUTE-time check must catch it, not just the282 gate-time check.283- `resources/matrix_w5.py` — **[coached]** interrupt a running turn (steer) then use the session284 again. Caught a real bug: the durable mount never gets re-established after a steer, breaking285 every subsequent turn on that session. Distinct from Mahmoud's own steer repro (that one is a286 turn-currency/heartbeat bug; this is a mount-lifecycle bug) — keep the two separate when287 triaging.288- `resources/matrix_w7.py` — **[coached]** an agent writes a workspace file and commits it via an289 `@ag.file` marker; asserts the approval manifest carries digest+bytes and the commit lands with290 the exact bytes. Caught a real bug (fixed, PR #5763-adjacent runner fix): the gate approved291 cleanly but execution always refused with `authorization_missing` — no file-marker commit could292 land. Re-verified PASS after the fix. Note: `DEFERRED_NOT_EXECUTED` on a queued second tool293 call is benign (another gate is already pending), not a real tool error — don't let it fail294 this cell. Naming `@ag.file` verbatim in the prompt is correct for this cell's backend-path295 purpose but means it cannot catch a model failing to find the marker syntax unprompted — see296 Tiers above.297- `resources/matrix_w1_daytona.py` — **[coached]** the commit round-trip on `sandbox=daytona`298 (every other cell runs local). Needs a funded provider vault key (see the file's own docstring299 for the exact vault-secret shape and the `provider_key` slug gotcha). PASS confirms300 `DaytonaWorkspaceReader` and the placeholder-secrets flow live.301- `resources/matrix_t8_saved_files.py` — **[coached]** T8: the Daytona remote agent mount (the302 durable `agent-files/` folder the playground file drawer writes into), which needs the ngrok303 tunnel the runner discovers at sandbox-acquire time. Writes a marker file into the mount via304 the mounts API (`POST /mounts/agents/sign` + `PUT /mounts/{id}/files`) before the run, then305 asserts the runner log line `remote agent mount active for artifact=<id>` appears, a tool306 output actually carries the real marker content (not hallucinated), and the commit lands with307 it. Verified PASS 3/3 runs after the tunnel-seat fix (2026-08-06); this line was absent on308 every attempt before that fix landed.309- `resources/matrix_t9_agent_tools.py` — **[coached]** T9: the runner restores the agent's own310 tools from `agent-files/.tools/` before a session (`agent-tools-setup.ts`). Plants a311 `setup.sh` and a `bin/qa-tool` through the mounts API, opens a fresh session, and asserts the312 first tool call is the exact probe and its output payload carries both planted tokens, plus a313 line `setup.sh` appended to `agent-files/.tools/runs.log`, read back through the mounts API314 with no model in the loop, plus the `agent_tools_setup` stage for THIS session in the runner315 log when `--runner-container` is given. `--sandbox local|daytona`, `--harness pi_core|claude`.316 Mandatory (via `path_triggers.py`) when the restore step or the sandbox image recipes change.317- `resources/matrix_w7_per_harness.py` — **[coached]** matrix_w7.py's exact scenario run318 identically on all three harnesses (claude, codex, pi_core), each classified PASS/FAIL/SKIP319 independently. Exists because W7 originally ran on Claude only, and that scenario-coverage gap320 is exactly what let the Codex approve-then-fail P0 (2026-08-06) ship: commit_revision + file321 marker + HITL approval on a harness this suite never exercised. claude uses subscription auth322 (no vault dependency); codex and pi_core need a funded OpenAI `provider_key` vault secret323 (mirrors cells X1/C3 in `qa_product.py`) and correctly SKIP with the exact reason when it's324 missing or ambiguous — a SKIP here is an untested harness, not a pass, and must be named as such325 in any release summary. `--only <harness>` runs a single leg without re-spending budget on the326 others. Verified PASS on claude (2026-08-06). codex/pi initially SKIPPED on this shared preview327 stack because its vault held zero — then, transiently, an ambiguous multiple — OpenAI candidates328 (other concurrent agents' activity on the same shared project); resolved 2026-08-06 by stocking329 one unambiguous OpenAI `provider_key` secret (no stale entries existed to remove). Re-verified330 PASS on codex after stocking the key (session 58ce3a58-8d04-40ac-99e4-c44eaa5d7b06).331- `resources/matrix_w7_daytona.py` — **[coached]** matrix_w7.py's exact scenario with332 `sandbox=daytona` instead of local. The local-only original W7 is exactly why this bug hid: the333 Daytona transport rejects NUL bytes in argv, which the `@ag.file` manifest walk was emitting, so334 no workspace-file commit could EVER land on Daytona, on any harness, until the 2026-08-06 fix335 (found during the same P0 triage, live-verified twice on codex+Daytona: sessions f3fa4335,336 f2f22056). This cell is the sandbox-axis regression guard, staying on the claude harness. Needs337 the same funded Anthropic vault key as `matrix_w1_daytona.py`. Verified PASS (2026-08-06).338- `resources/matrix_invariant_commit_auth_refusal.py` — **[coached scenario; the invariant itself339 is mechanism-level]** the generic invariant: no `tool_result` with empty output and340 `isError:false` may exist for a call whose runner log says `[commit-auth] refused` (the341 silent-blank-success class — the P0's actual failure shape, distinct from the scenario-coverage342 gap `matrix_w7_per_harness.py` addresses). The check itself343 (`qa_matrix_lib.check_no_blank_success_on_refusal`) reads the runner's own log line against the344 wire outcome and is meant to be reusable by any cell that exercises marker-carrying commits, not345 just this one. This cell's trigger is best-effort: run W7's flow, let the legitimate commit346 consume its authorization record, then REPLAY the byte-identical approval-carrying request347 (a duplicate submission) hoping to force a second, doomed `authorizeExecution` attempt. Verified348 2026-08-06: the replay did NOT reproduce a refusal (the runner treats the identical replayed349 history as already-resolved and answers conversationally instead of re-invoking the tool), so350 the cell correctly SKIPped rather than claiming a false pass on an invariant it never exercised.351 A reliable deterministic trigger (e.g. a genuine cold-resume stale-approval replay, or a crafted352 duplicate `toolCallId`) is an open follow-up; until then, treat any SKIP from this cell as "the353 invariant was not tested this run," never as green.354- `qa_matrix_lib.check_no_silent_turn(turns)` — **[mechanism-level invariant; no cell of its355 own]** no turn may come back completely bare: no text, no tool call, no approval gate, no file356 or data payload, and no error. That combination is a swallowed provider failure (ASD-EST100) —357 the model call is rejected, the error is dropped on the way back, and the turn is reported as a358 clean empty finish, so the user sees a blank bubble with no reason anywhere. It matters most in359 cells whose PASS depends on something NOT appearing (no error, no leak, no blank success): a360 turn that produced nothing satisfies those by doing nothing at all. Its content definition361 deliberately mirrors `content_parts_emitted` in the product's own Vercel egress362 (`sdks/python/agenta/sdk/agents/adapters/vercel/stream.py`) — reasoning does NOT count, so a363 turn that only thought is still a violation. Wired into `matrix_w7*.py`, `matrix_t8_saved_files.py`,364 `matrix_b1_builtin_find.py`, `matrix_invariant_commit_auth_refusal.py`,365 `matrix_l3_abandoned_approval.py`, `matrix_w3.py`, `matrix_w4.py` and `matrix_w5.py` as366 `... and not silent["violations"]`; add the same conjunct to any new cell. The one thing a cell367 must exclude itself is a turn it deliberately aborted or interrupted, which legitimately ends368 bare — `matrix_w5.py` shows the pattern (it checks only its post-interrupt turns).369- `resources/matrix_g1_guidance_discovery.py` — **[mechanism-blind]** does the platform guidance370 actually change what the model does? The trial prompt is Mahmoud's own verbatim phrasing from371 the live session that found the bug ("can you add gstack-autoplan skill to your skills (i saved372 it in your folder)") — no tool, operation, or marker syntax named. Before the guidance existed,373 this exact phrasing made a model copy the skill into its own harness-local skills folder and374 claim success without ever proposing a commit, 3/3 live (session b59cb549). Two parts per375 harness/model/sandbox leg: PROBE reads the rendered instructions file out of the workspace and376 asserts the fenced platform-guidance block and the skill-location sentence are really there;377 TRIALS run the prompt N times, PASS only when a commit_revision gate fires, is approved, and the378 STORED revision carries the skill (copying into the harness's own folder is a FAIL). Setup379 gotchas baked into `qa_matrix_lib.py` (`PI_CORE_HARNESS_KIND`, `PI_CORE_HAIKU_MODEL`): the Pi380 harness kind enum is `"pi_core"` (bare `"pi"` 500s), and `pi_core` rejects a bare `"haiku"` model381 id (needs `"claude-haiku-4-5"`); codex accepts its curated short alias (`"gpt-5.6-luna"`) bare.382 Both legs need `sandbox=daytona` + a vault key (codex: OpenAI; pi_core: the same Anthropic key383 `matrix_w1_daytona.py` documents) and SKIP with the exact reason when the credential is missing384 or ambiguous. `--only <leg>` runs a single leg without re-spending trial budget on the other.385 Neither leg has yet scored a clean 4/4 live, and both failure modes are real, reproducible model386 mechanics misses, not noise or infra flake — treat this cell's discovery rate as a genuine open387 quality question, not a settled pass:388 - **pi_core: 2/3** (2026-08-06). Trial 2's model ran `cp -r agent-files/gstack-autoplan389 .agenta-imports/` (copying the whole directory) then referenced a marker path that didn't390 match where the file landed; the engine correctly denied it fail-closed391 (`approved-content resolution failed ...: gstack-autoplan/SKILL.md does not exist under392 .agenta-imports/.; deny`). Not a product bug — the deny is doing its job — but evidence the393 model fumbles the exact copy-then-reference mechanics some of the time.394 - **codex: 2/3** (2026-08-06, re-verified after stocking a clean OpenAI vault key — the earlier395 SKIP was purely the missing/ambiguous credential, now resolved). Trial 1 didn't attempt the396 mechanism at all: zero gates approved, the model just replied "I can't add skills by simply397 placing files in the skills folder; skills must be enabled through the agent configuration"398 (session 4fa17164-3ada-4ef6-86b5-63bf1c64f10b) — it named the right concept but never called399 read_config/commit_revision to act on it. Trials 2 and 3 passed cleanly (sessions400 4c6a758b-39a3-4789-890e-06f3d68196b6, bf8a1283-667c-45c7-ab0d-b112e12106db).401402 Worth a call: whether the guidance text needs to be more directive (e.g. explicitly say "copy403 the FILE, not the directory" and "always attempt the tool calls, don't just describe the404 mechanism"), or whether ~2/3 is an acceptable bar for this feature's launch.405406### Builtin capability cells (`matrix_b*.py`) — does the harness's own tooling still work407408- `resources/matrix_b1_builtin_find.py` — **[coached, harness-mechanism test]** one native409 file-search call per harness (write three known marker files, then ask the model to locate410 them via its own search capability — never a manual directory listing), asserting the exact411 filenames come back in the TOOL OUTPUT payload, never the reply. Exists because nothing else in412 the gate ever exercised a harness builtin: verify-runner's overnight diagnosis found Pi's413 `find` builtin dead 52/52 across two benchmark runs (it shells out to the vendored `fd` binary414 with a flag that only exists from fd 9 onward; the runner image ships fd 8.6.0) — a total415 capability loss that sat invisible with nothing calling it. This cell closes that discoverability416 gap for the class, not just this one instance. **Open discrepancy, not yet reconciled**: this417 cell's own pi_core leg PASSED twice, live, with real filenames back418 (2026-08-07, sessions dd9c51ef-92d0-4780-855e-da6f48e07d9f and419 47f02ab1-61ec-4db8-a0d6-2d96e98b9188) — which does not match "52/52 failed". Either the break is420 conditional on a flag/option this cell's simple case never exercises, or something already421 changed; needs reconciling with verify-runner before Pi's `find` gets called either fixed or422 still broken. **Codex SKIPs by design**, not tested: its exec output doesn't land in the423 `tool-output-available` payload's `.output` field (the same quirk `qa_product.py`'s `j2_mount`424 already names and skips codex for), so this cell's evidence extraction cannot see codex's real425 results — a codex-shaped extraction is a follow-up. Claude PASSED cleanly on the two-turn426 version (session 4081e9ee-11c1-4f12-8245-6c390f90e9d8). Needed two EXPLICIT turns (write, then427 search) — one combined instruction left claude stopping after the write step without428 attempting the search at all.429430### The lifecycle cells (`matrix_l*.py`) — cold ↔ warm, and what survives each transition431432These four cover the session-lifecycle work: which config changes are applied to a RUNNING433sandbox and which tear it down, and what happens to a pending approval, a client tool and the434durable mount across each transition. They all assert the STORED turn ledger435(`POST /sessions/turns/query` → one `sandbox_id` per turn) or the STORED interaction rows436(`POST /sessions/interactions/query`), never the SSE echo — nothing about warm-versus-rebuilt437ever reaches the stream. An empty ledger FAILS a cell; missing evidence is not evidence.438439- `resources/matrix_l1_lifecycle_routes.py` — **MANDATORY. [mechanism-blind]** the routing matrix440 itself: for each kind of mid-conversation config change, assert the route the runner took. One441 sandbox id = applied in place, two = rebuilt. Blocks on six cases: no change must stay warm; an442 instructions edit, a permissions edit and a tool-catalog edit must escalate; and a443 same-connection model switch must stay warm on BOTH claude and pi_core. The pi_core model case444 (added 2026-08-29) is the standing trap for the wire-spelling bug class: the router once keyed445 its table on the bare "pi" literal while the wire carries "pi_core", every playground model446 switch silently rebuilt, and the claude-only case could not see it (#6364). This is the cell447 that would have caught the `cold1` rot described below.448- `resources/matrix_l2_approval_across_config_change.py` — **MANDATORY. [coached]** the killer449 combination: an approval answered while a config change rides along in the SAME request. It is450 the regression test for the applied-state bug (the pool used to stamp the INCOMING fingerprint451 on the approval-resume path, so the next turn continued warm on an environment running452 something else). Asserts the gated commit lands, the approval row ends `resolved`/`responded`,453 and — the real tell — the config change is not swallowed: it takes effect on the FOLLOWING454 turn, via a rebuild, in both the instructions and the permissions variant.455- `resources/matrix_l3_abandoned_approval.py` — **MANDATORY. [coached]** the user sends a new456 message instead of answering the card. Asserts the gated tool does NOT run (an unanswered457 approval is not consent), the row is swept to `cancelled` rather than left `pending`, and the458 session still works. `cancelled` vs `pending` is the loud-vs-silent distinction: a `pending` row459 is a card sitting on the page that no process is waiting on.460- `resources/matrix_l5_live_route_observed.py` — **MANDATORY. [mechanism-blind, with a control]**461 the other half of L1: an instructions edit made mid-conversation must actually be OBSERVED by462 the harness, not merely written to disk. Runs the same configuration on a fresh cold session as463 a control, so a failure isolates the runner rather than blaming the model; when the control also464 fails it reports INCONCLUSIVE instead of a confident wrong verdict. It asserts the edit, never465 the route, so it stays meaningful if the facet is ever made live again. **Failed 2026-08-06466 (claude/local) and now passes** — see the finding note below. Extended 2026-08-06 (overnight467 gate run) to all three harnesses in one invocation (`--only <harness>` for a single leg) —468 this MANDATORY blocker cell had only ever run on claude; codex and pi_core were a named gap.469 Verified PASS on all three the same night the harness matrix landed.470- `resources/matrix_l4_client_tool_lifecycle.py` — **nice-to-have. [coached]** the client-tool471 round trip, and the only cell that covers client tools at all. Asserts the browser's result472 reaches the model and the `client_tool` interaction is stored. It RECORDS rather than asserts473 the sandbox count, which is two today: a client-tool pause is deliberately not parkable474 (`"warm-hold": RESERVED, not built`, #5384), so every client-tool round trip currently costs a475 rebuild. If that number ever reads one, the warm hold landed and the docstring needs updating.476- `resources/matrix_n1_session_context.py` — **MANDATORY. [journey, with two controls]** the477 per-turn session facts on the path the PRODUCT uses. Renames a session between two turns and478 asks the agent for the name, asks the agent for its own display name, and posts a forged479 `meta.session_context` that the service must ignore. Every expected value carries a random480 token minted for the run and spoken nowhere in the conversation, so a transcript-derived481 answer cannot match; the second ask is the exact shape of #6661, because by then the FIRST482 name is in the transcript and the current one is only in the stored header. This is the one483 cell that asserts on model prose, and it does so because nothing else can: `turnContext` is a484 prompt string on the service-to-runner payload, the runner logs nothing, and it is485 deliberately kept out of `request.messages` and out of persisted input. Two controls keep a486 FAIL honest — an echo probe (a model that cannot repeat a literal token makes the cell report487 INCONCLUSIVE) and a read-back of the stored header after every rename. Mandatory (via488 `path_triggers.py`) when the SDK session-context module, the platform-prompt renderer, the489 agent handler, or the API's session-context resolver changes. **The gate could not see this490 class at all before v0.115.3**: #6661 shipped green because the API stamped the facts in its491 invoke prelude, which never runs for a playground turn, and no cell renamed a session492 mid-conversation or asserted on the facts.493494**Finding the lifecycle cells surfaced (2026-08-06, claude on local, reproduced 3×) — FIXED:** the495`workspaceFiles` live route rewrote the instruction file and advanced applied state, but the496running harness never re-read it. A warm session kept obeying the instructions it started with497while the pool reported the NEW fingerprint, so every later turn matched and continued warm and the498user's edit had no effect until something else evicted the session. A cold session with the499identical configuration obeyed it immediately, which is what isolated the runner. This was the500failure `desired-state.ts` refuses to allow for the `prompts` facet ("refreshing them and claiming501the model saw the change would be a lie") reappearing on the facet that WAS made live — and note502the direction: before the live route existed, an instructions edit forced a rebuild and therefore503took effect on the next turn, so it was a regression in what the user sees, not a speedup.504`matrix_l5_live_route_observed.py` is the repro.505506The fix withdrew the route: `workspaceFiles` now routes to `rebuild-sandbox` in the capability507table, and `refresh-workspace` left `LIVE_ACTION_KINDS` so restoring the table alone fails closed.508An instructions edit costs a sandbox again, which is what it cost before the optimisation. **L1's509`instructions` case therefore expects TWO sandbox ids, and L2's first variant expects two as510well** — if you are reading an old green from before 2026-08-06, those cells expected one. The511intended next shape is refresh THEN reopen the session, which needs the reopen to build its512session init from the incoming request first, and needs proving on L5 rather than asserting.513514**Why `cold1` changed (read before trusting an old green):** that tier used to force its eviction515by editing `instructions.agents_md`, which the lifecycle work briefly made a LIVE route — the tier516would have gone on passing while measuring warm reuse, and `park`'s "one sandbox id is meaningful"517argument rests on `cold1` reporting two on the same deployment. It now moves `harness.permissions`518(the `harnessSession` facet → `reopen-session`, deliberately not live) and ASSERTS two distinct519sandbox ids. It sta520521…(truncated)