Workflow 1.5: Experiment
Implement and deploy experiments from plan: $ARGUMENTS
Overview
This skill bridges Workflow 1 (idea discovery + method refinement) and Workflow 2 (auto review loop). It takes the experiment plan and turns it into running experiments with initial results.
Workflow 1 output: This skill: Workflow 2 input:
refine-logs/EXPERIMENT_PLAN.md → implement → LLM review → deploy → collect → initial results ready
refine-logs/EXPERIMENT_TRACKER.md code (cross-model) /run-experiment for /auto-iteration-loop
refine-logs/FINAL_PROPOSAL.md
Constants
- RESEARCH_DOMAIN = auto — Project domain tag (free-form, e.g.
mechanistic-interpretability, vision-encoders, rl-policy-eval). Consumed by Phase 1.5 only as a routing constraint to /mechanism-skills — see Phase 1.5 Step 2's domain: arg. When null or auto, Phase 1.5 infers from FINAL_PROPOSAL.md; on ambiguous inference, silently default to general and log [research-domain] inference ambiguous — defaulted to general (this fallback bypasses AUTO_PROCEED by design — see /auto's flag-table row for the canonical statement). To force a specific domain, pass it explicitly on the CLI. (Note: Phase 1.1 routes through /experiment-tips using its own symptom-level trigger table and does not consume this constant.)
- MECHANISM_ROUTING = auto — Phase 1.5 mechanism-family routing mode.
auto (default): invoke /mechanism-skills, write refine-logs/MECHANISM_ROUTING.md, present 2–3 candidates and let the caller pick (auto-select #1 when AUTO_PROCEED=true; otherwise block on the caller's AskUserQuestion). skip: assume routing already exists (or is not applicable) and proceed. not-applicable: explicitly mark behavioral-only proposal and skip without invoking. When called from /auto, the orchestrator's mini-prompt fills CHOSEN_FAMILY so this skill is re-entered with MECHANISM_ROUTING=skip.
- CHOSEN_FAMILY = none (dynamic — not in config; forwarded by /auto's orchestrator — from
MECHANISM=given (the user's chosen_mechanism captured by the claim stage), the AUTO_PROCEED=false family mini-prompt, or an explicit family: pin in task.md (cross-round Rule 2), after any settled-pin conflict is resolved) — When set, commits this family/submethod combo from MECHANISM_ROUTING.md before implementation (Phase 1.5 Mode B).
- CODE_REVIEW = true — external LLM reviewer checks experiment code before deployment. Catches logic bugs before wasting GPU hours. Set
false to skip.
- AUTO_DEPLOY = true — Automatically deploy experiments after implementation + review. Set
false to manually inspect code before deploying. Treated as a standing approval for the deploy step: when AUTO_DEPLOY=true, the Phase 4 deploy proceeds even if AUTO_PROCEED=false.
- AUTO_PROCEED = true — Whether the Phase 4 Experiment Gate may skip the UI prompt. When
true (default) and AUTO_DEPLOY=true, the gate proceeds silently. When false and AUTO_DEPLOY=false, the gate calls AskUserQuestion (approve / narrow-scope / abort) and blocks until the user answers. AUTO_DEPLOY=true overrides AUTO_PROCEED=false for this gate (standing approval). Forwarded from /auto.
- SANITY_FIRST = true — Run the sanity-stage experiment first (smallest, fastest) before launching the rest. Catches setup bugs early.
- MAX_PARALLEL_RUNS = 4 — Maximum number of experiments to deploy in parallel (limited by available GPUs). For Phase 4's queue dispatch path (Phase 4.B), this becomes
max_parallel: in the /experiment-queue manifest. For the direct dispatch path (Phase 4.A), it's the in-skill concurrency cap on /run-experiment calls.
- BATCH_DISPATCH =
auto — Phase 4 dispatch routing rule. auto (default): per the Phase 4.0 table — milestones with ≥ 10 runs, depends_on, grid expansions, or ≥ 3-seed × ≥ 3-config multi-seed sweeps go to /experiment-queue; smaller ad-hoc milestones go to /run-experiment. queue: force every milestone to /experiment-queue (use when you know the workload benefits from OOM retry + stale cleanup even at small sizes). direct: force every milestone to /run-experiment (use only when debugging the queue scheduler itself; emits a warning if any milestone would have triggered the queue rule under auto). Forwarded from /auto.
- BASE_REPO = null — GitHub repo URL to use as base codebase. When set, clone the repo first and implement experiments on top of it. When
null, write code from scratch or reuse existing project files.
- COMPACT = false — When
true, (1) read idea-stage/IDEA_CANDIDATES.md instead of full idea-stage/IDEA_REPORT.md if available, (2) append experiment results to EXPERIMENT_LOG.md after collection.
- RESUME = false — When
true, each phase checks if its primary artifact already exists non-empty and skips itself if so (see "Resume protocol" below). Useful for picking up after a crash. Default false = every phase always runs from scratch and overwrites prior artifacts. Resume never deletes pre-existing files.
- GPU_ID =
auto — GPU device(s) to use for sanity and full-suite runs. auto (default) inherits from environment / launcher. A single id (0) or comma-list (4,5,6,7) causes Phase 3 (sanity) and Phase 4 (deploy) to pass CUDA_VISIBLE_DEVICES=<GPU_ID> as the first positional argument to /run-experiment — the run-experiment skill then exports this env var before launching the experiment subprocess (do not treat it as a shell prefix; /run-experiment is a Skill invocation, not a shell command). Also record the effective CUDA_VISIBLE_DEVICES into each run's run.sh so reproductions land on the same devices. Override: — gpu-id: 4,5,6,7. When GPU_ID lists multiple devices and MAX_PARALLEL_RUNS > 1, partition devices across concurrent runs (e.g., GPU_ID=4,5,6,7 + 2 parallel → run A on 4,5, run B on 6,7); do not co-schedule two runs on the same device unless memory measurements confirm fit. Forwarded from /auto and from agents/experiment.md; /auto-verify follows the same convention for verify variants.
Standalone overrides: /auto-experiment "EXPERIMENT_PLAN.md" — compact: true, base-repo: https://github.com/org/project, research-domain: vision-encoders, mechanism-routing: auto.
Resource-Fidelity Harness (the reproduction combination)
Active only when the resource_fidelity: strict marker is present in the top metadata of refine-logs/FINAL_PROPOSAL.md / refine-logs/EXPERIMENT_PLAN.md — which /auto-claim stamps iff BEHAVIOR_SOURCE=given AND MECHANISM=given (the reproduction combination). There is no flag and no override: activation is purely marker-driven (only the reproduction combination stamps it; every other combination never does). When the marker is absent (any cost-aware combination), the existing cost-aware behavior is unchanged.
Scope: the main experiment only. The harness binds the main-experiment runs this skill (/auto-experiment) produces. It does not bind /auto-verify, whose deliberate model/dataset/method swaps are how it measures robustness — verify ignores the marker by design. When active, every phase of this skill obeys all five rules:
- Exact models. Instantiate the precise model id(s) / size(s) the plan specifies. Never substitute a smaller / cheaper / distilled / quantized variant to save compute or memory (quantization or reduced precision is allowed only if the plan /
task.md explicitly calls for it).
- Exact data. Use the full specified dataset(s) and
used_n / split. Never subset, down-sample, cap, or truncate the data to save time. A run that consumed less than the specified used_n is recorded failed, not done, and surfaced — never papered over with a subset note.
- No must-run skipped. Run every must-run experiment at full specified scale (seeds, configs, grid points). Cost (GPU-hours) is not grounds to skip or thin a must-run milestone.
- OOM → auto-scale-up across GPUs, never downscale. On OOM, resolve it only by science-neutral memory techniques — batch size ↓ (with gradient accumulation to preserve the effective batch), gradient checkpointing, sequence chunking, and automatically adding more GPUs to the run: query free GPUs (
nvidia-smi, memory.used < 500 MiB), add them and enable sharding, relaunch. If the script is single-GPU (model.cuda() / .to(device)), auto-convert its loading to device_map="auto" / FSDP / CPU-or-disk offload — and before trusting it on the full model, verify the conversion is wired correctly with a numeric-equivalence check on a fit-on-one-GPU proxy slice (single-GPU vs sharded outputs/loss must match within tolerance; sharding is a mathematically equivalent transform, so this step only catches a mis-wired edit, not the sharding itself). Keep adding free GPUs up to OOM_MAX_GPUS (default 4) while leaving MAX_PARALLEL_RUNS headroom for sibling runs. Only when the cap is reached (or no free GPU remains) and CPU/disk offload is also exhausted do you HALT and report — write Pipeline status = halted-at-experiment: strict-OOM-after-scaleup with the scale-up trace (started on N GPUs, auto-expanded to M, sharding mode, equivalence-check result, final memory shortfall) and next-round options (add GPUs / bigger machine / explicitly authorize quantization in task.md / explicitly down-scope the plan / switch to discovery). Never shrink the model or subset the data to make it fit — that is the one forbidden remedy, even in full-auto.
- Sanity is exempt but not a substitute. The
SANITY_FIRST smoke test may run at a tiny scale (its only job is catching setup bugs). The real runs that produce results must use the exact resources above; never report a sanity-scale run as a reproduction result.
Inputs
This skill expects one or more of:
refine-logs/EXPERIMENT_PLAN.md (best) — claim-driven experiment roadmap from /experiment-plan
refine-logs/EXPERIMENT_TRACKER.md — run-by-run execution table
refine-logs/FINAL_PROPOSAL.md — method description for implementation context
idea-stage/IDEA_CANDIDATES.md — compact idea summary (preferred when COMPACT: true) (fall back to ./IDEA_CANDIDATES.md if not found)
idea-stage/IDEA_REPORT.md — full brainstorm output (fall back to ./IDEA_REPORT.md if not found)
If none exist, ask the user what experiments to implement.
Workflow
Resume protocol (only when RESUME = true)
Skip entirely if RESUME = false (default). When true, each phase below begins with a skip-if-present check:
| Phase |
Primary artifact (skip key) |
Notes |
| 1 |
refine-logs/EXPERIMENT_PLAN.md already loaded into memory (no separate file) |
Always cheap; just re-parse. |
| 1.1 |
refine-logs/EXPERIMENT_TIPS.md with committed: true |
Skip routing if already committed; matched tips must still be re-read into the current context each Phase 2 invocation, since context doesn't survive across calls. |
| 1.25 |
refine-logs/EXPERIMENT_RESULTS.md with phenomenon_status: set to a terminal value (not-established, or inconclusive after the retry budget) |
The M0 gate already ran and ended the stage — do not re-run M0 or proceed to mechanism. Re-emit the same terminal phenomenon_status in the return so /auto re-applies the early exit. If phenomenon_status is established / conditional (or M0 partially ran), fall through and let Phase 1.5+ resume normally. Only applies when the plan has a milestone with kind: phenomenon-validation (BEHAVIOR_SOURCE ∈ {given-validation, discovery}); otherwise this phase doesn't exist. |
| 1.5 |
refine-logs/MECHANISM_ROUTING.md with committed: true (or routing: not-applicable) and a ## Plan reconciliation section whose reconciliation_status is set (ok/escalate/n/a) |
Skip routing entirely only if both hold. If committed: true but reconciliation is missing (a prior run crashed between commit and Step 7), do not skip — resume at Step 7 to complete reconciliation before Phase 2. If reconciliation_status: escalate, HALT (do not proceed to build). |
| 2 |
runs/<run_id>_<short_purpose>/ directory non-empty for every run referenced by the plan |
Per-run granularity: skip code generation for runs whose directory already has code. Path convention matches the "Output Directory Naming" section below. |
| 2.5 |
reviewer-approved marker in run directory (e.g., .code_review_passed) |
Per-run. |
| 3 |
sanity run's results captured in EXPERIMENT_TRACKER.md |
Skip sanity if its row is present. |
| 4 |
per-run status in EXPERIMENT_TRACKER.md marked done |
Per-run: don't redeploy completed runs. |
| 5 |
refine-logs/EXPERIMENT_RESULTS.md exists non-empty AND EXPERIMENT_TRACKER.md has every plan run marked terminal |
Skip collection when everything is already aggregated. |
| 5.5 / 5.6 / 6 |
always run (cheap rewrites / handoff steps) |
— |
Log every skip as [resume] phase <N> skipped — <reason>. Resume never deletes pre-existing files. To force a phase to re-run, delete its primary artifact (or, for Phase 4, the run directory).
Phase 1: Parse the Experiment Plan
Read EXPERIMENT_PLAN.md and extract:
- Run order and milestones — which experiments run first (sanity → baseline → main → ablation → polish)
- For each experiment block:
- Dataset / split / task
- Compared systems and variants
- Metrics to compute
- Setup details (backbone, hyperparameters, seeds)
- Success criterion
- Priority (MUST-RUN vs NICE-TO-HAVE)
- Compute budget — total estimated GPU-hours
- Method details from
FINAL_PROPOSAL.md — what exactly to implement
- Mechanism routing hints (under
mechanism_strategy:, written by the claim stage from cross-round memory) — the chosen direction; any families_already_settled: [<families>] (families already confirmed/refuted for this behavior+direction, to be excluded at routing — Phase 1.5); and any explicit family: pin. The list is absent (or omitted) when there is nothing to exclude — round 1, a new phenomenon, or a non-mechanism plan.
Present a brief summary:
📋 Experiment plan loaded:
- Milestones: [N] (sanity → baseline → main → ablation)
- Must-run experiments: [N]
- Nice-to-have: [N]
- Estimated GPU-hours: [X]
Proceeding to implementation.
Data Rules — load and check now (data design). Must load skills/data-rule/SKILL.md and validate the plan's data against its four rules (provenance / splits / labels / sample-size floor). Loaded once here, it governs all data use through the rest of the workflow — unconditional, not the Phase 1.1 symptom routing. Surface any violation before proceeding.
Phase 1.1: Experiment-Tips Routing
Before any code is written, route the parsed plan through /experiment-tips — the routing entry point under skills/experiment-tips/. Tips encode hard-won conventions that prevent silent reproducibility / overclaim failure modes. Catching them at plan-implementation time is free; catching them after a full deploy wastes GPU hours.
Routing flow:
Read the artifacts that drive the routing: refine-logs/EXPERIMENT_PLAN.md, refine-logs/FINAL_PROPOSAL.md. Extract: task description, datasets, models, declared n_pairs / n_examples, intervention sites (block / layer / token / residue), steering parameter names (α, dose, magnitude, scale, coefficient), and any adapter-based fine-tune milestone (LoRA / QLoRA / PEFT).
Invoke /experiment-tips as the routing entry point. It matches the plan against the symptom-level trigger table in skills/experiment-tips/SKILL.md and returns a list of tip folders to load.
For each matched tip, load skills/experiment-tips/<tip>/SKILL.md in full into working memory. Hard requirement: the routing previews in experiment-tips/SKILL.md are deliberately thin — every implementation detail lives in the tip's SKILL.md, not in the routing file. Acting on the preview alone is forbidden.
Write refine-logs/EXPERIMENT_TIPS.md with the routing decision:
# Experiment Tips Routing
<!-- Metadata block (parsed by /auto orchestrator resume check). -->
committed: true
matched_tips:
- <tip-folder-1>
- <tip-folder-2>
## Matches
1. **<tip-folder>** — <one-line trigger that fired>
- convention to adopt: <one-line summary from the tip's own SKILL.md — not from the routing preview>
2. **<tip-folder>** — <one-line trigger that fired>
- convention to adopt: <one-line summary from the tip's own SKILL.md — not from the routing preview>
## No-match log
<if a borderline trigger didn't fire, note it here for audit>
If no tip matches, write a stub EXPERIMENT_TIPS.md containing committed: true and matched_tips: [] plus a one-line note explaining why none matched (e.g., "behavioral-only proposal, no representation interventions"). Always write the file — it's the audit anchor for Phase 2 to confirm tips were considered.
Hard requirement: before Phase 1.25, the experiment agent's transcript must show an actual /experiment-tips invocation. Bypassing the skill or fabricating refine-logs/EXPERIMENT_TIPS.md by hand is forbidden. If the transcript shows no invocation, or the file is absent or hand-written, re-run Phase 1.1.
Hard requirement: every matched_tips entry in EXPERIMENT_TIPS.md and its convention to adopt must be adopted in all downstream experiments.
Phase 1.25: Phenomenon-Validation Gate (BEHAVIOR_SOURCE ∈ {given-validation, discovery})
Runs only when EXPERIMENT_PLAN.md contains a milestone carrying the machine marker kind: phenomenon-validation (conventionally titled M0) — i.e. the claim stage ran with BEHAVIOR_SOURCE ∈ {given-validation, discovery}. Detect by the kind: phenomenon-validation field, never by the milestone's title (the title may be phrased or localized freely; only the field is the stable contract). If no milestone has that marker (a BEHAVIOR_SOURCE=given run, or any non-phenomenon plan), skip this phase entirely and set phenomenon_status = n/a. The principle: claim only assumed the phenomenon exists and proposed a mechanism for it — this gate is where the assumption is actually tested, before any mechanism compute is spent.
Run M0 first and alone at the plan's specified used_n, then decide a four-state phenomenon_status. Concretely:
- Implement only M0's code (the Phase 2 implementation step, scoped to M0's milestone). If
CODE_REVIEW=true, M0's code gets one cross-model review pass (Phase 2.5) like any other run. M0's scale is set by the plan's used_n — the agent may not subset, cap, or down-sample M0's data to save compute. An M0 run that consumed less than the planned used_n is recorded failed, not done (Phase 5's used_n rule).
- Deploy M0 via
/run-experiment (one run). This creates M0's runs/<run_id>_*/ directory and flips its EXPERIMENT_TRACKER.md row to done/failed through the normal mechanism — so the later Phase 4 deploy automatically skips re-running M0 (its per-run skip key is already satisfied), and resume is consistent. Do not hand-roll separate bookkeeping.
- Forward
GPU_ID to M0's run exactly as Phase 4 would.
Then decide the verdict:
Lightweight integrity check on M0 before trusting its verdict: invoke /experiment-audit scoped to M0 (the same audit verify's Phase 2 uses). This separates "the phenomenon is genuinely absent" from "the M0 test itself was broken/underpowered".
Assign phenomenon_status:
established — M0 audit is clean AND the behavior reproduces per its pass criterion (paraphrase/seed/decoding robust, confounds controlled, adequate n, trivial-explanation ruled out). → Proceed to Phase 1.5 and the mechanism milestones (M1…Mn) normally.
conditional — M0 audit clean, but the behavior holds only under a subset of conditions. → Runtime-scope; do not edit the plan (EXPERIMENT_PLAN.md is owned by claim Phase 4.5). Run the mechanism milestones with inputs restricted to the condition subset where the phenomenon holds, record phenomenon_status: conditional + the boundary in EXPERIMENT_RESULTS.md, and tag the claim conditional — holds under <X> so /auto carries it into the ledger. Formally narrowing the claim text and re-planning is the iteration loop's type-③ job, not this stage. AUTO_PROCEED: true auto-continues with the narrowed scope; false → AskUserQuestion (continue with narrowed scope / terminate) and block.
inconclusive — M0's integrity audit FAILs; the phenomenon is untested, not disproven. → Do not terminate. Diagnose, fix at the lowest sufficient level, and re-run M0. Log every fix under M0's block in EXPERIMENT_RESULTS.md (what changed, old → new, audit rationale).
- Script bug (crash, wrong path, tensor-shape error, eval bug) →
auto-debug M0 (≤ 3 attempts), then re-run.
- Run-level methodology (under-sampled within plan limits, seed not fixed, decoding drift) → adjust the M0 run invocation only; leave the plan alone.
- Plan-level methodology (wrong metric,
used_n too low, mis-specified coefficient / threshold / hyperparameter, inadequate dataset size or split) → most commonly under-tuned knobs (finetune lr / LoRA rank / epochs / batch, steering α / target layer, thresholds, decoding temperature). Re-invoke /experiment-tips explicitly for hyperparameter / coefficient tuning — not a general re-audit — passing M0's realized evidence (loss / effect size / refusal rate / seed variance) and routing to the matching sub-skill: experiment-tips/finetune-hyperparameter-sweep for fine-tune knobs, experiment-tips/steering-coefficient-tuning for α / β / dose of any additive intervention (steering / CAA / DAS / SAE / ROME), experiment-tips/steering-block-selection for the target layer / site / window. Apply the returned search range / scan grid by editing EXPERIMENT_PLAN.md in place, confined to the implicated field(s); do not touch the claim, phenomenon description, or milestone graph.
not-established — M0 audit clean AND behavior does not reproduce at adequate power. → Before settling a terminal negative, treat the null as a candidate hyperparameter / coefficient mis-setting: re-invoke /experiment-tips for tuning with M0's realized evidence, using the same sub-skill routing as the plan-level branch above. If a matching sub-skill returns a concrete scan or knob change, apply it under the inconclusive branch's script / run / plan rules and re-run M0; otherwise fall through to step 3.
Tuning budget and terminal handling (shared by inconclusive and not-established). Total Phase 1.25 iterations (initial deploy + tuning retries) ≤ 5; the budget covers all tuning attempts across knobs, not one per knob. Terminate when any of the following holds:
(a) the 5-iteration budget is exhausted and the verdict is still not-established / inconclusive,
(b) no matching tuning sub-skill returns an actionable change, or
(c) you diagnose a hyperparameter / coefficient as a likely contributor to the null but the remaining budget cannot cover the required search space.
On termination: write terminal refine-logs/EXPERIMENT_RESULTS.md with phenomenon_status: not-established or inconclusive in the top metadata (use inconclusive, not not-established, when a narrow or absent coefficient sweep is a suspected cause), M0's evidence framed as a negative result (or as unresolved, for inconclusive), and every retry (knob, old → new value, resulting verdict) logged under M0's block. Append a mandatory entry to claims_ledger.json's open_items[] (rendered into CLAIMS_LEDGER.md's Open Items section) of the form: "Phase 1.25 stopped after <n> automated tuning iterations on <milestone id> (knobs tried: <knob>=<v0>→<v1>, …); phenomenon remained <status>, and hyperparameters / coefficients are a likely contributor. Recommend the user manually tune from the recorded scan bounds before re-invoking /auto." This open-item entry is required whenever tuning is a suspected cause — trigger (c) covers the case where you stop before exhausting the budget because the automated sweep is not the right instrument. Set phenomenon_status in the Phase 6 return so /auto skips verify + iteration and ends as ended-phenomenon-not-established or ended-phenomenon-inconclusive. AUTO_PROCEED: true auto-terminates; false → AskUserQuestion (terminate — accept the result and write the report (recommended) / re-run M0 — I will adjust the test/plan first) and block. Do not offer "narrow scope" — that is a conditional-only action.
🚨 Phase 1.25 hard constraints (non-negotiable):
- ~20 GPU-hours are pre-allocated. Time/GPU cost is not a valid reason to skip a sweep or retry — only the terminal triggers below (5-iteration budget / no actionable sub-skill / trigger (c)) may stop iteration.
- Terminal
not-established / inconclusive skips all downstream mechanism milestones (M1…Mn). You must exhaust matching tuning sub-skills and retries within the 5-attempt budget before settling on either verdict.
Mechanism milestones declare depends_on: [M0], so even outside this gate the queue will not launch them before M0 completes; the gate adds the verdict-based branch (run status alone is not enough — a clean M0 that shows no effect must stop the pipeline, which depends_on cannot express).
Phase 1.5: Mechanism-Family Routing
A committed mechanism family is a hard precondition for any Phase 2 code that touches an internal object (layer, head, neuron, SAE feature, weight, input feature).
CHOSEN_FAMILY takes precedence over MECHANISM_ROUTING. Whenever CHOSEN_FAMILY is set to a real family (anything other than unset / none / not-applicable), go straight to the routing flow's Mode-B commit (Step 6) regardless of the MECHANISM_ROUTING value — Step 6 commits CHOSEN_FAMILY (and builds the routing scaffold around it when MECHANISM_ROUTING.md is absent or does not list it). This is what makes all three direct-commit sources work: the first-class MECHANISM=given path and the family: pin (no Mode A ran → file may be absent → Step 6 builds it), and the AUTO_PROCEED=false post-mini-prompt build call (Mode A already wrote the candidates → Step 6 flips committed). A bare MECHANISM_ROUTING=skip therefore never silently proceeds without a committed family when CHOSEN_FAMILY is set. CHOSEN_FAMILY = not-applicable is the behavioral-only sentinel (a behavioral-only proposal, incl. the MECHANISM=given behavioral-only reproduction): treat it exactly like MECHANISM_ROUTING = not-applicable — write the routing: not-applicable stub, run no mechanism milestone — regardless of the MECHANISM_ROUTING value.
When CHOSEN_FAMILY is unset, behavior is gated by MECHANISM_ROUTING:
MECHANISM_ROUTING = skip — assume refine-logs/MECHANISM_ROUTING.md already exists (or is intentionally absent for a behavioral proposal). Read it if present, otherwise log [routing] skipped — no manifest, proceeding without routing and continue to Phase 2.
MECHANISM_ROUTING = not-applicable — write a stub refine-logs/MECHANISM_ROUTING.md containing routing: not-applicable and a one-line justification pulled from FINAL_PROPOSAL.md, then continue.
MECHANISM_ROUTING = auto (default) — run the routing flow below.
Routing flow (auto mode):
Use the Phase 1 parse already in memory to identify routing inputs: chosen claim(s), the internal objects the method targets (layer / head / neuron / SAE feature / weight / input feature), the read sites, any mechanism families already named in FINAL_PROPOSAL.md, and the cross-round routing hints from EXPERIMENT_PLAN.md (item 5) — the families_already_settled: [<families>] avoid-set (Rule 1) and any family: pin. No new file reads needed at this step.
Invoke /mechanism-skills via the Skill tool — this is a hard requirement, not a description of behavior. The manifest must be derived from the catalog this skill loads at routing time, not from prior training knowledge. You are not permitted to write MECHANISM_ROUTING.md until you have actually read, in this turn, via the Skill tool:
skills/mechanism-skills/SKILL.md (the routing entry point listing all eleven families)
- The
SKILL.md of every family you intend to list as a candidate
- The
SKILL.md of every submethod underneath those families
If any of these files is unread, re-invoke /mechanism-skills before generating candidates — do not write the manifest from memory.
Pass RESEARCH_DOMAIN as the routing constraint:
- If set to a specific tag (e.g.,
mechanistic-interpretability), pass it as — domain: <RESEARCH_DOMAIN> so candidates are restricted to the matching families.
- If
auto, do not pass — domain:; ask /mechanism-skills to infer from FINAL_PROPOSAL.md. On ambiguous inference, fall back to domain: general silently (same rule as the Constants section).
Produce 2–3 candidate family/submethod combinations from the catalog. Each candidate must use the canonical family name as it appears in skills/mechanism-skills/SKILL.md — one of: Causal Attribution, Circuit Discovery, Probing, Magnitude Analysis, Gradient Detection, Feature Dictionary Learning, Representation and Parameter Analysis, Vocabulary Projection, SHAP, Neural Feature Learning, Multi-Modal.The chosen_family and candidate_paths metadata fields must use canonical catalog names.
Cross-round avoid-set (Rule 1): if families_already_settled (item 5) lists any families, exclude them from the candidate list — they were already confirmed/refuted for this behavior+direction, so re-routing to them redoes settled work. A family left inconclusive is not settled and may still be proposed (ideally with a refined submethod). Note the exclusion in ## Rationale. If excluding leaves no viable catalog candidate for the direction, do not silently fall back to a settled family — surface it (the direction may itself be exhausted; a signal to pick a different direction or behavior next round).
For each candidate, record:
- Canonical
family/submethod
path of the form skills/mechanism-skills/<family>/<submethod>/SKILL.md. Verify each path exists on disk (Glob or ls) before writing the manifest; if any candidate path is missing, that candidate is a hallucination — re-invoke /mechanism-skills and reground.
- Planned screen → decode → verify → recover composition
- Cost notes (GPU-hours / wall-clock estimate)
- One-line rationale tied to a specific claim
- Effective domain (the resolved value from step 2)
Pure downstream analysis steps — matrix factorization, low-rank decomposition, linear algebra on already-collected effect vectors, etc. — are not mechanism families and must not occupy the chosen_family slot. They belong in the composition plan as post-processing.
Write refine-logs/MECHANISM_ROUTING.md with: inputs read, candidate list (mark #1 as recommended), composition plan with cost notes, and the routing rationale. The file MUST contain an explicit committed: line in its top metadata block (see template below) — Mode A writes committed: false, Mode B (CHOSEN_FAMILY set) overwrites it to committed: true. The /auto orchestrator's four-branch resume protocol keys directly off this string, so omitting the line will break resume.
Required template for refine-logs/MECHANISM_ROUTING.md:
# Mechanism Routing
<!-- Metadata block (parsed by /auto orchestrator resume check). -->
committed: <false|true>
chosen_family: <canonical family/submethod | none>
chosen_idea_title: <title | n/a>
effective_domain: <resolved domain>
candidate_paths:
- skills/mechanism-skills/<family>/<submethod>/SKILL.md
- skills/mechanism-skills/<family>/<submethod>/SKILL.md
## Candidates
1. **[recommended]** <canonical family/submethod> — <rationale>
- path: skills/mechanism-skills/<family>/<submethod>/SKILL.md
2. <canonical family/submethod> — <rationale>
- path: skills/mechanism-skills/<family>/<submethod>/SKILL.md
3. <canonical family/submethod> — <rationale>
- path: skills/mechanism-skills/<family>/<submethod>/SKILL.md
## Composition plan
<screen → decode → verify → recover, with cost notes. Downstream analysis steps (decomposition, probing aggregation, …) belong here, not in the candidate slot.>
## Plan reconciliation
<!-- Written by Step 7 once a family is committed. One row per method_sensitive field declared on the intervention milestone(s). -->
- n_pairs: plan=<X> → <matches | re-bound <Y> — <why the committed submethod needs it> | conflict — <why this cannot be satisfied without changing the plan's scientific intent>>
- sites: plan=<...> → <matches | re-bound <...> — <why> | conflict — <why>>
- metric: plan=<...> → <matches | re-bound <...> — <why> | conflict — <why>>
- gpu_hours: plan~<X> → revised ~<Y> — <what in the committed submethod's compute profile drives the change; may go up OR down, e.g. a gradient-based approximation adds a backward pass but avoids per-site forward passes, while exhaustive activation patching scales with the number of patched sites>
reconciliation_status: <ok | escalate | n/a> <!-- n/a when the milestone(s) declared no method_sensitive fields (incl. the reproduction combo, which pins them exact); escalate iff any field is `conflict`; else ok -->
## Rationale
<why #1 is recommended; if aligned_with_tagging: no, explain why the prior was over-ruled>
For the routing: not-applicable short-circuit, replace the metadata block with routing: not-applicable and committed: true (a behavioral-only proposal is "committed" to having no mechanism family) and keep a one-line justification under ## Rationale.
If CHOSEN_FAMILY is unset — branch on AUTO_PROCEED:
true (default): auto-select the [recommended] candidate, flip committed: true, continue into Phase 2 in the same call. No mini-prompt. Log [routing] auto-selected family=<name> submethod=<name>.
false: write committed: false, return to caller (Mode A complete). Caller (/auto orchestrator or standalone CLI) prompts the user then re-enters with CHOSEN_FAMILY set.
If CHOSEN_FAMILY is set (Mode B — build) — commit that family. Two sub-cases:
MECHANISM_ROUTING.md exists and already lists CHOSEN_FAMILY as a candidate (the AUTO_PROCEED=false mini-prompt path, where Mode A ran first) — locate that candidate, overwrite the metadata line committed: false to committed: true (and set chosen_family: to it), and proceed to Phase 2.
MECHANISM_ROUTING.md is missing, or exists but does not list CHOSEN_FAMILY (the first-class MECHANISM=given path and the cross-round family: pin path — no Mode A ran, so the file may be absent or its auto-candidates may not include the user's pick) — do not auto-select #1 and do not silently drop the user's choice. Run steps 1–4 to build the routing scaffold (composition plan, candidate_paths, rationale) for CHOSEN_FAMILY specifically: load /mechanism-skills and map CHOSEN_FAMILY to the nearest canonical catalog family/submethod — the user may have written a free-text method name (e.g. activation patching → Causal Attribution / activation patching; SAE → Feature Dictionary Learning / sparse-autoencoder), so resolve it semantically against the catalog rather than requiring an exact string. If it resolves to a catalog entry, write MECHANISM_ROUTING.md with the canonical name as the sole committed candidate (committed: true, chosen_family: <canonical name>). If it plausibly maps to no catalog family/submethod, HALT with [routing] CHOSEN_FAMILY="<x>" maps to no /mechanism-skills family/submethod — name a supported mechanism method in task.md. This is the "commit the named family directly" behavior the CHOSEN_FAMILY sources note describes.
In both sub-cases log [routing] committed family=<name> submethod=<name> and proceed to Step 7 (Plan reconciliation) then Phase 2.
CHOSEN_FAMILY sources. CHOSEN_FAMILY is forwarded by /auto's orchestrator from one of three places: (1) MECHANISM=given — the user named the mechanism method/family in task.md and the claim stage stamped it as chosen_mechanism in FINAL_PROPOSAL.md / EXPERIMENT_PLAN.md (this is the first-class mechanism-given path: no routing, no mini-prompt); (2) the AUTO_PROCEED=false family mini-prompt (MECHANISM=discovery); (3) an explicit family: pin in task.md (the cross-round Rule-2 path under MECHANISM=discovery — see /auto → Global Exploration Memory). Mode B treats all three identically: commit the named family (per Step 6 — auto-selection of a different family never happens; the scaffold is built around CHOSEN_FAMILY, not re-routed away from it). A pinned family that is itself in families_already_settled is fine — the orchestrator already confirmed the re-run with the user (honor-pin) before forwarding — so commit it rather than blocking.
Plan reconciliation (runs whenever a family becomes committed — the auto-select in step 5 or the Mode B commit in step 6; skip for routing: not-applicable). Now that a concrete submethod is locked, reconcile it against the fields the claim stage could only estimate before the method was known. Read the just-loaded submethod SKILL.md (already in context from step 2) for its real requirements, then for every method_sensitive field declared on the intervention milestone(s) in EXPERIMENT_PLAN.md, compare the plan's value to what this submethod needs and write the verdict into the ## Plan reconciliation section of MECHANISM_ROUTING.md:
matches — the plan's value is fine for this submethod. Nothing else to do.
re-bound <new value> — <why> — the submethod needs a different value that still serves the milestone's scientific intent (e.g. attribution-patching's gradient estimate needs more n_pairs for a stable estimate; a method needs a different read site; the GPU-hours estimate moves up or down because the committed submethod's compute profile differs from the generic pre-routing estimate). Record the new value here; do not edit EXPERIMENT_PLAN.md (the plan stays the claim stage's audit reference — the realized value is captured downstream as planned-vs-actual in Phase 5). Phase 2 implements to the re-bound value; the Phase 4 gate displays the re-bound GPU-hours.
conflict — <why> — the submethod cannot satisfy the field without changing the plan's scientific intent (e.g. it structurally cannot measure the planned metric, or requires sites the plan explicitly excluded). This is a plan defect, not a routing re-bind, and is out of this stage's authority (this skill never rewrites the claim-authored plan). Set reconciliation_status: escalate and stop before Phase 2 (do not build). The orchestrator surfaces this as a Round-End Decision (ended-needs-decision, never a crash halt and never an auto-rewrite of the plan), and the plan owner (the user) repairs the conflicting field in EXPERIMENT_PLAN.md — or picks a fitting submethod, or re-scopes the claim — and re-runs. Do not silently swap the metric/sites to make the method fit.
**Resource-F
…(truncated)
1---2name: auto-experiment3description: Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, routes mechanism family inline (Phase 1.5), implements experiment code, deploys to GPU, and collects initial results. Use when user says "implement experiments", "experiment", "deploy the plan", or has an experiment plan ready to execute.4---56# Workflow 1.5: Experiment78Implement and deploy experiments from plan: **$ARGUMENTS**910## Overview1112This skill bridges Workflow 1 (idea discovery + method refinement) and Workflow 2 (auto review loop). It takes the experiment plan and turns it into running experiments with initial results.1314```15Workflow 1 output: This skill: Workflow 2 input:16refine-logs/EXPERIMENT_PLAN.md → implement → LLM review → deploy → collect → initial results ready17refine-logs/EXPERIMENT_TRACKER.md code (cross-model) /run-experiment for /auto-iteration-loop18refine-logs/FINAL_PROPOSAL.md19```2021## Constants2223- **RESEARCH_DOMAIN = auto** — Project domain tag (free-form, e.g. `mechanistic-interpretability`, `vision-encoders`, `rl-policy-eval`). **Consumed by Phase 1.5 only** as a routing constraint to `/mechanism-skills` — see Phase 1.5 Step 2's `domain:` arg. When `null` or `auto`, Phase 1.5 infers from `FINAL_PROPOSAL.md`; on ambiguous inference, silently default to `general` and log `[research-domain] inference ambiguous — defaulted to general` (this fallback bypasses `AUTO_PROCEED` by design — see `/auto`'s flag-table row for the canonical statement). To force a specific domain, pass it explicitly on the CLI. (Note: Phase 1.1 routes through `/experiment-tips` using its own symptom-level trigger table and does **not** consume this constant.)24- **MECHANISM_ROUTING = auto** — Phase 1.5 mechanism-family routing mode. `auto` (default): invoke `/mechanism-skills`, write `refine-logs/MECHANISM_ROUTING.md`, present 2–3 candidates and let the caller pick (auto-select #1 when `AUTO_PROCEED=true`; otherwise block on the caller's `AskUserQuestion`). `skip`: assume routing already exists (or is not applicable) and proceed. `not-applicable`: explicitly mark behavioral-only proposal and skip without invoking. When called from `/auto`, the orchestrator's mini-prompt fills `CHOSEN_FAMILY` so this skill is re-entered with `MECHANISM_ROUTING=skip`.25- **CHOSEN_FAMILY = none** *(dynamic — not in config; forwarded by /auto's orchestrator — from `MECHANISM=given` (the user's `chosen_mechanism` captured by the claim stage), the `AUTO_PROCEED=false` family mini-prompt, or an explicit `family:` pin in `task.md` (cross-round Rule 2), after any settled-pin conflict is resolved)* — When set, commits this family/submethod combo from `MECHANISM_ROUTING.md` before implementation (Phase 1.5 Mode B).26- **CODE_REVIEW = true** — external LLM reviewer checks experiment code before deployment. Catches logic bugs before wasting GPU hours. Set `false` to skip.27- **AUTO_DEPLOY = true** — Automatically deploy experiments after implementation + review. Set `false` to manually inspect code before deploying. Treated as a *standing approval* for the deploy step: when `AUTO_DEPLOY=true`, the Phase 4 deploy proceeds even if `AUTO_PROCEED=false`.28- **AUTO_PROCEED = true** — Whether the Phase 4 Experiment Gate may skip the UI prompt. When `true` (default) and `AUTO_DEPLOY=true`, the gate proceeds silently. When `false` and `AUTO_DEPLOY=false`, the gate calls `AskUserQuestion` (approve / narrow-scope / abort) and blocks until the user answers. `AUTO_DEPLOY=true` overrides `AUTO_PROCEED=false` for this gate (standing approval). Forwarded from `/auto`.29- **SANITY_FIRST = true** — Run the sanity-stage experiment first (smallest, fastest) before launching the rest. Catches setup bugs early.30- **MAX_PARALLEL_RUNS = 4** — Maximum number of experiments to deploy in parallel (limited by available GPUs). For Phase 4's queue dispatch path (Phase 4.B), this becomes `max_parallel:` in the `/experiment-queue` manifest. For the direct dispatch path (Phase 4.A), it's the in-skill concurrency cap on `/run-experiment` calls.31- **BATCH_DISPATCH = `auto`** — Phase 4 dispatch routing rule. `auto` (default): per the Phase 4.0 table — milestones with ≥ 10 runs, `depends_on`, grid expansions, or ≥ 3-seed × ≥ 3-config multi-seed sweeps go to `/experiment-queue`; smaller ad-hoc milestones go to `/run-experiment`. `queue`: force every milestone to `/experiment-queue` (use when you know the workload benefits from OOM retry + stale cleanup even at small sizes). `direct`: force every milestone to `/run-experiment` (use only when debugging the queue scheduler itself; emits a warning if any milestone would have triggered the queue rule under `auto`). Forwarded from `/auto`.32- **BASE_REPO = null** — GitHub repo URL to use as base codebase. When set, clone the repo first and implement experiments on top of it. When `null`, write code from scratch or reuse existing project files.33- **COMPACT = false** — When `true`, (1) read `idea-stage/IDEA_CANDIDATES.md` instead of full `idea-stage/IDEA_REPORT.md` if available, (2) append experiment results to `EXPERIMENT_LOG.md` after collection.34- **RESUME = false** — When `true`, each phase checks if its primary artifact already exists non-empty and skips itself if so (see "Resume protocol" below). Useful for picking up after a crash. Default `false` = every phase always runs from scratch and overwrites prior artifacts. Resume never deletes pre-existing files.35- **GPU_ID = `auto`** — GPU device(s) to use for sanity and full-suite runs. `auto` (default) inherits from environment / launcher. A single id (`0`) or comma-list (`4,5,6,7`) causes Phase 3 (sanity) and Phase 4 (deploy) to **pass `CUDA_VISIBLE_DEVICES=<GPU_ID>` as the first positional argument to `/run-experiment`** — the run-experiment skill then exports this env var before launching the experiment subprocess (do not treat it as a shell prefix; `/run-experiment` is a Skill invocation, not a shell command). Also record the effective `CUDA_VISIBLE_DEVICES` into each run's `run.sh` so reproductions land on the same devices. Override: `— gpu-id: 4,5,6,7`. When `GPU_ID` lists multiple devices and `MAX_PARALLEL_RUNS > 1`, partition devices across concurrent runs (e.g., `GPU_ID=4,5,6,7` + 2 parallel → run A on `4,5`, run B on `6,7`); do not co-schedule two runs on the same device unless memory measurements confirm fit. Forwarded from `/auto` and from `agents/experiment.md`; `/auto-verify` follows the same convention for verify variants.3637> Standalone overrides: `/auto-experiment "EXPERIMENT_PLAN.md" — compact: true, base-repo: https://github.com/org/project, research-domain: vision-encoders, mechanism-routing: auto`.3839## Resource-Fidelity Harness (the reproduction combination)4041**Active only when the `resource_fidelity: strict` marker is present** in the top metadata of `refine-logs/FINAL_PROPOSAL.md` / `refine-logs/EXPERIMENT_PLAN.md` — which `/auto-claim` stamps **iff `BEHAVIOR_SOURCE=given` AND `MECHANISM=given`** (the reproduction combination). There is no flag and no override: activation is purely marker-driven (only the reproduction combination stamps it; every other combination never does). When the marker is absent (any cost-aware combination), the existing cost-aware behavior is unchanged.4243**Scope: the main experiment only.** The harness binds the main-experiment runs this skill (`/auto-experiment`) produces. It does **not** bind `/auto-verify`, whose deliberate model/dataset/method swaps are how it measures robustness — verify ignores the marker by design. When active, every phase of *this* skill obeys all five rules:44451. **Exact models.** Instantiate the precise model id(s) / size(s) the plan specifies. Never substitute a smaller / cheaper / distilled / quantized variant to save compute or memory (quantization or reduced precision is allowed only if the plan / `task.md` explicitly calls for it).462. **Exact data.** Use the full specified dataset(s) and `used_n` / split. Never subset, down-sample, cap, or truncate the data to save time. A run that consumed less than the specified `used_n` is recorded **failed**, not `done`, and surfaced — never papered over with a subset note.473. **No must-run skipped.** Run every must-run experiment at full specified scale (seeds, configs, grid points). Cost (GPU-hours) is not grounds to skip or thin a must-run milestone.484. **OOM → auto-scale-up across GPUs, never downscale.** On OOM, resolve it **only** by science-neutral memory techniques — batch size ↓ (with gradient accumulation to preserve the effective batch), gradient checkpointing, sequence chunking, and **automatically adding more GPUs to the run**: query free GPUs (`nvidia-smi`, `memory.used < 500 MiB`), add them and enable sharding, relaunch. If the script is single-GPU (`model.cuda()` / `.to(device)`), auto-convert its loading to `device_map="auto"` / FSDP / CPU-or-disk offload — and before trusting it on the full model, verify the conversion is wired correctly with a **numeric-equivalence check on a fit-on-one-GPU proxy slice** (single-GPU vs sharded outputs/loss must match within tolerance; sharding is a mathematically equivalent transform, so this step only catches a mis-wired edit, not the sharding itself). Keep adding free GPUs up to `OOM_MAX_GPUS` (default 4) while leaving `MAX_PARALLEL_RUNS` headroom for sibling runs. **Only** when the cap is reached (or no free GPU remains) **and** CPU/disk offload is also exhausted do you **HALT and report** — write `Pipeline status = halted-at-experiment: strict-OOM-after-scaleup` with the scale-up trace (started on N GPUs, auto-expanded to M, sharding mode, equivalence-check result, final memory shortfall) and next-round options (add GPUs / bigger machine / explicitly authorize quantization in `task.md` / explicitly down-scope the plan / switch to discovery). **Never** shrink the model or subset the data to make it fit — that is the one forbidden remedy, even in full-auto.495. **Sanity is exempt but not a substitute.** The `SANITY_FIRST` smoke test may run at a tiny scale (its only job is catching setup bugs). The real runs that produce results must use the exact resources above; never report a sanity-scale run as a reproduction result.5051## Inputs5253This skill expects one or more of:54551. **`refine-logs/EXPERIMENT_PLAN.md`** (best) — claim-driven experiment roadmap from `/experiment-plan`562. **`refine-logs/EXPERIMENT_TRACKER.md`** — run-by-run execution table573. **`refine-logs/FINAL_PROPOSAL.md`** — method description for implementation context584. **`idea-stage/IDEA_CANDIDATES.md`** — compact idea summary (preferred when `COMPACT: true`) *(fall back to `./IDEA_CANDIDATES.md` if not found)*595. **`idea-stage/IDEA_REPORT.md`** — full brainstorm output *(fall back to `./IDEA_REPORT.md` if not found)*6061If none exist, ask the user what experiments to implement.6263## Workflow6465### Resume protocol (only when `RESUME = true`)6667Skip entirely if `RESUME = false` (default). When `true`, each phase below begins with a **skip-if-present** check:6869| Phase | Primary artifact (skip key) | Notes |70|---|---|---|71| 1 | `refine-logs/EXPERIMENT_PLAN.md` already loaded into memory (no separate file) | Always cheap; just re-parse. |72| 1.1 | `refine-logs/EXPERIMENT_TIPS.md` with `committed: true` | Skip routing if already committed; matched tips must still be re-read into the current context each Phase 2 invocation, since context doesn't survive across calls. |73| 1.25 | `refine-logs/EXPERIMENT_RESULTS.md` with `phenomenon_status:` set to a **terminal** value (`not-established`, or `inconclusive` after the retry budget) | The M0 gate already ran and ended the stage — do **not** re-run M0 or proceed to mechanism. Re-emit the same terminal `phenomenon_status` in the return so `/auto` re-applies the early exit. If `phenomenon_status` is `established` / `conditional` (or M0 partially ran), fall through and let Phase 1.5+ resume normally. Only applies when the plan has a milestone with `kind: phenomenon-validation` (`BEHAVIOR_SOURCE ∈ {given-validation, discovery}`); otherwise this phase doesn't exist. |74| 1.5 | `refine-logs/MECHANISM_ROUTING.md` with `committed: true` (or `routing: not-applicable`) **and** a `## Plan reconciliation` section whose `reconciliation_status` is set (`ok`/`escalate`/`n/a`) | Skip routing entirely only if both hold. If `committed: true` but reconciliation is missing (a prior run crashed between commit and Step 7), do **not** skip — resume at Step 7 to complete reconciliation before Phase 2. If `reconciliation_status: escalate`, HALT (do not proceed to build). |75| 2 | `runs/<run_id>_<short_purpose>/` directory non-empty for every run referenced by the plan | Per-run granularity: skip code generation for runs whose directory already has code. Path convention matches the "Output Directory Naming" section below. |76| 2.5 | reviewer-approved marker in run directory (e.g., `.code_review_passed`) | Per-run. |77| 3 | sanity run's results captured in `EXPERIMENT_TRACKER.md` | Skip sanity if its row is present. |78| 4 | per-run status in `EXPERIMENT_TRACKER.md` marked `done` | Per-run: don't redeploy completed runs. |79| 5 | `refine-logs/EXPERIMENT_RESULTS.md` exists non-empty AND `EXPERIMENT_TRACKER.md` has every plan run marked terminal | Skip collection when everything is already aggregated. |80| 5.5 / 5.6 / 6 | always run (cheap rewrites / handoff steps) | — |8182Log every skip as `[resume] phase <N> skipped — <reason>`. Resume never deletes pre-existing files. To force a phase to re-run, delete its primary artifact (or, for Phase 4, the run directory).8384### Phase 1: Parse the Experiment Plan8586Read `EXPERIMENT_PLAN.md` and extract:87881. **Run order and milestones** — which experiments run first (sanity → baseline → main → ablation → polish)892. **For each experiment block:**90 - Dataset / split / task91 - Compared systems and variants92 - Metrics to compute93 - Setup details (backbone, hyperparameters, seeds)94 - Success criterion95 - Priority (MUST-RUN vs NICE-TO-HAVE)963. **Compute budget** — total estimated GPU-hours974. **Method details** from `FINAL_PROPOSAL.md` — what exactly to implement985. **Mechanism routing hints** (under `mechanism_strategy:`, written by the claim stage from cross-round memory) — the chosen `direction`; any `families_already_settled: [<families>]` (families already `confirmed`/`refuted` for this behavior+direction, to be **excluded** at routing — Phase 1.5); and any explicit `family:` pin. The list is absent (or omitted) when there is nothing to exclude — round 1, a new phenomenon, or a non-mechanism plan.99100Present a brief summary:101102```103📋 Experiment plan loaded:104- Milestones: [N] (sanity → baseline → main → ablation)105- Must-run experiments: [N]106- Nice-to-have: [N]107- Estimated GPU-hours: [X]108109Proceeding to implementation.110```111112**Data Rules — load and check now (data design).** Must load `skills/data-rule/SKILL.md` and validate the plan's data against its four rules (provenance / splits / labels / sample-size floor). Loaded once here, it governs all data use through the rest of the workflow — unconditional, not the Phase 1.1 symptom routing. Surface any violation before proceeding.113114### Phase 1.1: Experiment-Tips Routing 115116Before any code is written, route the parsed plan through `/experiment-tips` — the routing entry point under `skills/experiment-tips/`. Tips encode hard-won conventions that prevent silent reproducibility / overclaim failure modes. Catching them at plan-implementation time is free; catching them after a full deploy wastes GPU hours.117118**Routing flow:**1191201. **Read** the artifacts that drive the routing: `refine-logs/EXPERIMENT_PLAN.md`, `refine-logs/FINAL_PROPOSAL.md`. Extract: task description, datasets, models, declared `n_pairs` / `n_examples`, intervention sites (block / layer / token / residue), steering parameter names (`α`, `dose`, `magnitude`, `scale`, `coefficient`), and any adapter-based fine-tune milestone (LoRA / QLoRA / PEFT).1211222. **Invoke `/experiment-tips`** as the routing entry point. It matches the plan against the symptom-level trigger table in `skills/experiment-tips/SKILL.md` and returns a list of tip folders to load.1231243. **For each matched tip**, load `skills/experiment-tips/<tip>/SKILL.md` in full into working memory. **Hard requirement:** the routing previews in `experiment-tips/SKILL.md` are deliberately thin — every implementation detail lives in the tip's `SKILL.md`, not in the routing file. Acting on the preview alone is forbidden.1251264. **Write `refine-logs/EXPERIMENT_TIPS.md`** with the routing decision:127128 ```markdown129 # Experiment Tips Routing130131 <!-- Metadata block (parsed by /auto orchestrator resume check). -->132 committed: true133 matched_tips:134 - <tip-folder-1>135 - <tip-folder-2>136137 ## Matches138139 1. **<tip-folder>** — <one-line trigger that fired>140 - convention to adopt: <one-line summary from the tip's own SKILL.md — not from the routing preview>141 2. **<tip-folder>** — <one-line trigger that fired>142 - convention to adopt: <one-line summary from the tip's own SKILL.md — not from the routing preview>143144 ## No-match log145 <if a borderline trigger didn't fire, note it here for audit>146 ```147148 If no tip matches, write a stub `EXPERIMENT_TIPS.md` containing `committed: true` and `matched_tips: []` plus a one-line note explaining why none matched (e.g., "behavioral-only proposal, no representation interventions"). Always write the file — it's the audit anchor for Phase 2 to confirm tips were considered.149150**Hard requirement**: before Phase 1.25, the experiment agent's transcript must show an actual `/experiment-tips` invocation. Bypassing the skill or fabricating `refine-logs/EXPERIMENT_TIPS.md` by hand is forbidden. If the transcript shows no invocation, or the file is absent or hand-written, re-run Phase 1.1.151152**Hard requirement**: every `matched_tips` entry in `EXPERIMENT_TIPS.md` and its `convention to adopt` must be adopted in all downstream experiments.153154### Phase 1.25: Phenomenon-Validation Gate (`BEHAVIOR_SOURCE ∈ {given-validation, discovery}`)155156**Runs only when `EXPERIMENT_PLAN.md` contains a milestone carrying the machine marker `kind: phenomenon-validation`** (conventionally titled M0) — i.e. the claim stage ran with `BEHAVIOR_SOURCE ∈ {given-validation, discovery}`. **Detect by the `kind: phenomenon-validation` field, never by the milestone's title** (the title may be phrased or localized freely; only the field is the stable contract). If no milestone has that marker (a `BEHAVIOR_SOURCE=given` run, or any non-phenomenon plan), skip this phase entirely and set `phenomenon_status = n/a`. The principle: claim only *assumed* the phenomenon exists and proposed a mechanism for it — this gate is where the assumption is actually tested, **before** any mechanism compute is spent.157158159Run M0 **first and alone** at the plan's specified `used_n`, then decide a **four-state `phenomenon_status`**. Concretely:160161- **Implement** only M0's code (the Phase 2 implementation step, scoped to M0's milestone). If `CODE_REVIEW=true`, M0's code gets one cross-model review pass (Phase 2.5) like any other run. M0's scale is set by the plan's `used_n` — the agent may not subset, cap, or down-sample M0's data to save compute. An M0 run that consumed less than the planned `used_n` is recorded **failed**, not `done` (Phase 5's `used_n` rule).162- **Deploy** M0 via `/run-experiment` (one run). This creates M0's `runs/<run_id>_*/` directory and flips its `EXPERIMENT_TRACKER.md` row to `done`/`failed` through the normal mechanism — so the later Phase 4 deploy **automatically skips re-running M0** (its per-run skip key is already satisfied), and resume is consistent. Do not hand-roll separate bookkeeping.163- Forward `GPU_ID` to M0's run exactly as Phase 4 would.164165Then decide the verdict:1661671. **Lightweight integrity check** on M0 before trusting its verdict: invoke `/experiment-audit` scoped to M0 (the same audit verify's Phase 2 uses). This separates "the phenomenon is genuinely absent" from "the M0 test itself was broken/underpowered".1682. Assign `phenomenon_status`:169 - **`established`** — M0 audit is clean AND the behavior reproduces per its pass criterion (paraphrase/seed/decoding robust, confounds controlled, adequate n, trivial-explanation ruled out). → **Proceed** to Phase 1.5 and the mechanism milestones (M1…Mn) normally.170 - **`conditional`** — M0 audit clean, but the behavior holds only under a subset of conditions. → **Runtime-scope; do not edit the plan** (`EXPERIMENT_PLAN.md` is owned by claim Phase 4.5). Run the mechanism milestones with inputs restricted to the condition subset where the phenomenon holds, record `phenomenon_status: conditional` + the boundary in `EXPERIMENT_RESULTS.md`, and tag the claim `conditional — holds under <X>` so `/auto` carries it into the ledger. Formally narrowing the claim text and re-planning is the iteration loop's type-③ job, not this stage. **`AUTO_PROCEED`**: `true` auto-continues with the narrowed scope; `false` → `AskUserQuestion` (`continue with narrowed scope` / `terminate`) and block.171 - **`inconclusive`** — M0's integrity audit FAILs; the phenomenon is *untested*, not disproven. → Do not terminate. Diagnose, fix at the lowest sufficient level, and re-run M0. Log every fix under M0's block in `EXPERIMENT_RESULTS.md` (what changed, old → new, audit rationale).172 - **Script bug** (crash, wrong path, tensor-shape error, eval bug) → `auto-debug` M0 (≤ 3 attempts), then re-run.173 - **Run-level methodology** (under-sampled within plan limits, seed not fixed, decoding drift) → adjust the M0 **run invocation** only; leave the plan alone.174 - **Plan-level methodology** (wrong metric, `used_n` too low, mis-specified coefficient / threshold / hyperparameter, inadequate dataset size or split) → most commonly under-tuned knobs (finetune lr / LoRA rank / epochs / batch, steering α / target layer, thresholds, decoding temperature). **Re-invoke `/experiment-tips` explicitly for hyperparameter / coefficient tuning** — not a general re-audit — passing M0's realized evidence (loss / effect size / refusal rate / seed variance) and routing to the matching sub-skill: `experiment-tips/finetune-hyperparameter-sweep` for fine-tune knobs, `experiment-tips/steering-coefficient-tuning` for α / β / dose of any additive intervention (steering / CAA / DAS / SAE / ROME), `experiment-tips/steering-block-selection` for the target layer / site / window. Apply the returned search range / scan grid by editing `EXPERIMENT_PLAN.md` in place, confined to the implicated field(s); do not touch the claim, phenomenon description, or milestone graph.175176 - **`not-established`** — M0 audit clean AND behavior does **not** reproduce at adequate power. → Before settling a terminal negative, treat the null as a candidate hyperparameter / coefficient mis-setting: **re-invoke `/experiment-tips` for tuning** with M0's realized evidence, using the same sub-skill routing as the **plan-level branch** above. If a matching sub-skill returns a concrete scan or knob change, apply it under the `inconclusive` branch's script / run / plan rules and re-run M0; otherwise fall through to step 3.1771781793. **Tuning budget and terminal handling** (shared by `inconclusive` and `not-established`). Total Phase 1.25 iterations (initial deploy + tuning retries) ≤ **5**; the budget covers all tuning attempts across knobs, not one per knob. Terminate when any of the following holds:180 (a) the 5-iteration budget is exhausted and the verdict is still `not-established` / `inconclusive`,181 (b) no matching tuning sub-skill returns an actionable change, or182 (c) you diagnose a hyperparameter / coefficient as a likely contributor to the null but the remaining budget cannot cover the required search space.183184 On termination: write terminal `refine-logs/EXPERIMENT_RESULTS.md` with `phenomenon_status: not-established` or `inconclusive` in the top metadata (use `inconclusive`, not `not-established`, when a narrow or absent coefficient sweep is a suspected cause), M0's evidence framed as a negative result (or as unresolved, for `inconclusive`), and every retry (**knob, old → new value, resulting verdict**) logged under M0's block. **Append a mandatory entry to `claims_ledger.json`'s `open_items[]`** (rendered into `CLAIMS_LEDGER.md`'s Open Items section) of the form: *"Phase 1.25 stopped after `<n>` automated tuning iterations on `<milestone id>` (knobs tried: `<knob>=<v0>→<v1>`, …); phenomenon remained `<status>`, and hyperparameters / coefficients are a likely contributor. Recommend the user manually tune from the recorded scan bounds before re-invoking `/auto`."* This open-item entry is **required whenever tuning is a suspected cause** — trigger (c) covers the case where you stop before exhausting the budget because the automated sweep is not the right instrument. Set `phenomenon_status` in the Phase 6 return so `/auto` skips verify + iteration and ends as `ended-phenomenon-not-established` or `ended-phenomenon-inconclusive`. **`AUTO_PROCEED`**: `true` auto-terminates; `false` → `AskUserQuestion` (`terminate — accept the result and write the report` (recommended) / `re-run M0 — I will adjust the test/plan first`) and block. Do **not** offer "narrow scope" — that is a `conditional`-only action.185186> 🚨 **Phase 1.25 hard constraints (non-negotiable):**187> 1. **~20 GPU-hours are pre-allocated.** Time/GPU cost is **not** a valid reason to skip a sweep or retry — only the terminal triggers below (5-iteration budget / no actionable sub-skill / trigger (c)) may stop iteration.188> 2. **Terminal `not-established` / `inconclusive` skips all downstream mechanism milestones (M1…Mn).** You **must** exhaust matching tuning sub-skills and retries within the 5-attempt budget before settling on either verdict.189190Mechanism milestones declare `depends_on: [M0]`, so even outside this gate the queue will not launch them before M0 completes; the gate adds the *verdict-based* branch (run status alone is not enough — a clean M0 that shows no effect must stop the pipeline, which `depends_on` cannot express).191192### Phase 1.5: Mechanism-Family Routing193194A committed mechanism family is a **hard precondition** for any Phase 2 code that touches an internal object (layer, head, neuron, SAE feature, weight, input feature).195196**`CHOSEN_FAMILY` takes precedence over `MECHANISM_ROUTING`.** Whenever `CHOSEN_FAMILY` is set to a real family (anything other than unset / `none` / `not-applicable`), go straight to the routing flow's **Mode-B commit (Step 6) regardless of the `MECHANISM_ROUTING` value** — Step 6 commits `CHOSEN_FAMILY` (and builds the routing scaffold around it when `MECHANISM_ROUTING.md` is absent or does not list it). This is what makes all three direct-commit sources work: the first-class `MECHANISM=given` path and the `family:` pin (no Mode A ran → file may be absent → Step 6 builds it), and the `AUTO_PROCEED=false` post-mini-prompt build call (Mode A already wrote the candidates → Step 6 flips `committed`). A bare `MECHANISM_ROUTING=skip` therefore **never** silently proceeds without a committed family when `CHOSEN_FAMILY` is set. **`CHOSEN_FAMILY = not-applicable`** is the behavioral-only sentinel (a behavioral-only proposal, incl. the `MECHANISM=given` behavioral-only reproduction): treat it exactly like `MECHANISM_ROUTING = not-applicable` — write the `routing: not-applicable` stub, run no mechanism milestone — regardless of the `MECHANISM_ROUTING` value.197198When `CHOSEN_FAMILY` is **unset**, behavior is gated by `MECHANISM_ROUTING`:199200- **`MECHANISM_ROUTING = skip`** — assume `refine-logs/MECHANISM_ROUTING.md` already exists (or is intentionally absent for a behavioral proposal). Read it if present, otherwise log `[routing] skipped — no manifest, proceeding without routing` and continue to Phase 2.201- **`MECHANISM_ROUTING = not-applicable`** — write a stub `refine-logs/MECHANISM_ROUTING.md` containing `routing: not-applicable` and a one-line justification pulled from `FINAL_PROPOSAL.md`, then continue.202- **`MECHANISM_ROUTING = auto`** (default) — run the routing flow below.203204**Routing flow (auto mode):**2052061. **Use the Phase 1 parse already in memory** to identify routing inputs: chosen claim(s), the internal objects the method targets (layer / head / neuron / SAE feature / weight / input feature), the read sites, any mechanism families already named in `FINAL_PROPOSAL.md`, and the **cross-round routing hints** from `EXPERIMENT_PLAN.md` (item 5) — the `families_already_settled: [<families>]` avoid-set (Rule 1) and any `family:` pin. No new file reads needed at this step.2072082. **Invoke `/mechanism-skills` via the Skill tool — this is a hard requirement, not a description of behavior.** The manifest must be derived from the catalog this skill loads at routing time, not from prior training knowledge. You are not permitted to write `MECHANISM_ROUTING.md` until you have actually read, in this turn, via the Skill tool:209 - `skills/mechanism-skills/SKILL.md` (the routing entry point listing all eleven families)210 - The `SKILL.md` of every family you intend to list as a candidate211 - The `SKILL.md` of every submethod underneath those families212213 If any of these files is unread, re-invoke `/mechanism-skills` before generating candidates — do not write the manifest from memory.214215 Pass `RESEARCH_DOMAIN` as the routing constraint:216 - If set to a specific tag (e.g., `mechanistic-interpretability`), pass it as `— domain: <RESEARCH_DOMAIN>` so candidates are restricted to the matching families.217 - If `auto`, do not pass `— domain:`; ask `/mechanism-skills` to infer from `FINAL_PROPOSAL.md`. On ambiguous inference, fall back to `domain: general` silently (same rule as the Constants section).2182193. **Produce 2–3 candidate family/submethod combinations from the catalog.** Each candidate must use the **canonical family name** as it appears in `skills/mechanism-skills/SKILL.md` — one of: Causal Attribution, Circuit Discovery, Probing, Magnitude Analysis, Gradient Detection, Feature Dictionary Learning, Representation and Parameter Analysis, Vocabulary Projection, SHAP, Neural Feature Learning, Multi-Modal.The `chosen_family` and `candidate_paths` metadata fields must use canonical catalog names.220221 **Cross-round avoid-set (Rule 1):** if `families_already_settled` (item 5) lists any families, **exclude** them from the candidate list — they were already `confirmed`/`refuted` for this behavior+direction, so re-routing to them redoes settled work. A family left `inconclusive` is *not* settled and **may** still be proposed (ideally with a refined submethod). Note the exclusion in `## Rationale`. If excluding leaves no viable catalog candidate for the direction, do not silently fall back to a settled family — surface it (the direction may itself be exhausted; a signal to pick a different direction or behavior next round).222223 For each candidate, record:224 - Canonical `family/submethod`225 - `path` of the form `skills/mechanism-skills/<family>/<submethod>/SKILL.md`. Verify each path exists on disk (Glob or `ls`) before writing the manifest; if any candidate path is missing, that candidate is a hallucination — re-invoke `/mechanism-skills` and reground.226 - Planned screen → decode → verify → recover composition227 - Cost notes (GPU-hours / wall-clock estimate)228 - One-line rationale tied to a specific claim229 - Effective domain (the resolved value from step 2)230231 Pure downstream analysis steps — matrix factorization, low-rank decomposition, linear algebra on already-collected effect vectors, etc. — are **not** mechanism families and must not occupy the `chosen_family` slot. They belong in the composition plan as post-processing.2322334. **Write `refine-logs/MECHANISM_ROUTING.md`** with: inputs read, candidate list (mark #1 as recommended), composition plan with cost notes, and the routing rationale. **The file MUST contain an explicit `committed:` line in its top metadata block** (see template below) — Mode A writes `committed: false`, Mode B (`CHOSEN_FAMILY` set) overwrites it to `committed: true`. The `/auto` orchestrator's four-branch resume protocol keys directly off this string, so omitting the line will break resume.234235 **Required template for `refine-logs/MECHANISM_ROUTING.md`:**236237 ```markdown238 # Mechanism Routing239240 <!-- Metadata block (parsed by /auto orchestrator resume check). -->241 committed: <false|true>242 chosen_family: <canonical family/submethod | none>243 chosen_idea_title: <title | n/a>244 effective_domain: <resolved domain>245 candidate_paths:246 - skills/mechanism-skills/<family>/<submethod>/SKILL.md247 - skills/mechanism-skills/<family>/<submethod>/SKILL.md248249 ## Candidates250251 1. **[recommended]** <canonical family/submethod> — <rationale>252 - path: skills/mechanism-skills/<family>/<submethod>/SKILL.md253 2. <canonical family/submethod> — <rationale>254 - path: skills/mechanism-skills/<family>/<submethod>/SKILL.md255 3. <canonical family/submethod> — <rationale>256 - path: skills/mechanism-skills/<family>/<submethod>/SKILL.md257258 ## Composition plan259 <screen → decode → verify → recover, with cost notes. Downstream analysis steps (decomposition, probing aggregation, …) belong here, not in the candidate slot.>260261 ## Plan reconciliation262 <!-- Written by Step 7 once a family is committed. One row per method_sensitive field declared on the intervention milestone(s). -->263 - n_pairs: plan=<X> → <matches | re-bound <Y> — <why the committed submethod needs it> | conflict — <why this cannot be satisfied without changing the plan's scientific intent>>264 - sites: plan=<...> → <matches | re-bound <...> — <why> | conflict — <why>>265 - metric: plan=<...> → <matches | re-bound <...> — <why> | conflict — <why>>266 - gpu_hours: plan~<X> → revised ~<Y> — <what in the committed submethod's compute profile drives the change; may go up OR down, e.g. a gradient-based approximation adds a backward pass but avoids per-site forward passes, while exhaustive activation patching scales with the number of patched sites>267 reconciliation_status: <ok | escalate | n/a> <!-- n/a when the milestone(s) declared no method_sensitive fields (incl. the reproduction combo, which pins them exact); escalate iff any field is `conflict`; else ok -->268269 ## Rationale270 <why #1 is recommended; if aligned_with_tagging: no, explain why the prior was over-ruled>271 ```272273 For the `routing: not-applicable` short-circuit, replace the metadata block with `routing: not-applicable` and `committed: true` (a behavioral-only proposal is "committed" to having no mechanism family) and keep a one-line justification under `## Rationale`.2742755. **If `CHOSEN_FAMILY` is unset** — branch on `AUTO_PROCEED`:276 - `true` (default): auto-select the `[recommended]` candidate, flip `committed: true`, continue into Phase 2 in the same call. No mini-prompt. Log `[routing] auto-selected family=<name> submethod=<name>`.277 - `false`: write `committed: false`, return to caller (Mode A complete). Caller (`/auto` orchestrator or standalone CLI) prompts the user then re-enters with `CHOSEN_FAMILY` set.2782796. **If `CHOSEN_FAMILY` is set (Mode B — build)** — commit that family. Two sub-cases:280 - **`MECHANISM_ROUTING.md` exists and already lists `CHOSEN_FAMILY` as a candidate** (the `AUTO_PROCEED=false` mini-prompt path, where Mode A ran first) — locate that candidate, **overwrite the metadata line `committed: false` to `committed: true`** (and set `chosen_family:` to it), and proceed to Phase 2.281 - **`MECHANISM_ROUTING.md` is missing, or exists but does not list `CHOSEN_FAMILY`** (the first-class `MECHANISM=given` path and the cross-round `family:` pin path — no Mode A ran, so the file may be absent or its auto-candidates may not include the user's pick) — **do not auto-select #1 and do not silently drop the user's choice.** Run steps 1–4 to build the routing scaffold (composition plan, `candidate_paths`, rationale) for `CHOSEN_FAMILY` specifically: **load `/mechanism-skills` and map `CHOSEN_FAMILY` to the nearest canonical catalog family/submethod** — the user may have written a free-text method name (e.g. `activation patching` → `Causal Attribution / activation patching`; `SAE` → `Feature Dictionary Learning / sparse-autoencoder`), so resolve it semantically against the catalog rather than requiring an exact string. If it resolves to a catalog entry, write `MECHANISM_ROUTING.md` with the **canonical** name as the sole committed candidate (`committed: true`, `chosen_family: <canonical name>`). If it plausibly maps to **no** catalog family/submethod, HALT with `[routing] CHOSEN_FAMILY="<x>" maps to no /mechanism-skills family/submethod — name a supported mechanism method in task.md`. This is the "commit the named family directly" behavior the `CHOSEN_FAMILY` sources note describes.282283 In both sub-cases log `[routing] committed family=<name> submethod=<name>` and proceed to Step 7 (Plan reconciliation) then Phase 2.284285 **`CHOSEN_FAMILY` sources.** `CHOSEN_FAMILY` is forwarded by `/auto`'s orchestrator from one of three places: **(1) `MECHANISM=given`** — the user named the mechanism method/family in `task.md` and the claim stage stamped it as `chosen_mechanism` in `FINAL_PROPOSAL.md` / `EXPERIMENT_PLAN.md` (this is the first-class mechanism-given path: no routing, no mini-prompt); **(2)** the `AUTO_PROCEED=false` family mini-prompt (`MECHANISM=discovery`); **(3)** an explicit `family:` pin in `task.md` (the cross-round Rule-2 path under `MECHANISM=discovery` — see `/auto` → Global Exploration Memory). Mode B treats all three identically: **commit the named family** (per Step 6 — auto-selection of a *different* family never happens; the scaffold is built around `CHOSEN_FAMILY`, not re-routed away from it). A pinned family that is itself in `families_already_settled` is **fine** — the orchestrator already confirmed the re-run with the user (`honor-pin`) before forwarding — so commit it rather than blocking.2862877. **Plan reconciliation (runs whenever a family becomes committed — the auto-select in step 5 or the Mode B commit in step 6; skip for `routing: not-applicable`).** Now that a concrete submethod is locked, reconcile it against the fields the claim stage could only estimate before the method was known. Read the just-loaded submethod `SKILL.md` (already in context from step 2) for its real requirements, then for **every `method_sensitive` field** declared on the intervention milestone(s) in `EXPERIMENT_PLAN.md`, compare the plan's value to what this submethod needs and write the verdict into the `## Plan reconciliation` section of `MECHANISM_ROUTING.md`:288 - **`matches`** — the plan's value is fine for this submethod. Nothing else to do.289 - **`re-bound <new value> — <why>`** — the submethod needs a different value that still serves the milestone's *scientific intent* (e.g. attribution-patching's gradient estimate needs more `n_pairs` for a stable estimate; a method needs a different read `site`; the GPU-hours estimate moves up or down because the committed submethod's compute profile differs from the generic pre-routing estimate). Record the new value here; **do not edit `EXPERIMENT_PLAN.md`** (the plan stays the claim stage's audit reference — the realized value is captured downstream as planned-vs-actual in Phase 5). Phase 2 implements to the re-bound value; the Phase 4 gate displays the re-bound GPU-hours.290 - **`conflict — <why>`** — the submethod **cannot** satisfy the field without changing the plan's scientific intent (e.g. it structurally cannot measure the planned `metric`, or requires `sites` the plan explicitly excluded). This is a **plan defect, not a routing re-bind**, and is out of this stage's authority (this skill never rewrites the claim-authored plan). Set `reconciliation_status: escalate` and **stop before Phase 2** (do not build). The orchestrator surfaces this as a **Round-End Decision** (`ended-needs-decision`, never a crash halt and never an auto-rewrite of the plan), and the plan owner (the user) repairs the conflicting field in `EXPERIMENT_PLAN.md` — or picks a fitting submethod, or re-scopes the claim — and re-runs. Do **not** silently swap the metric/sites to make the method fit.291292 **Resource-F293294…(truncated)