Craft Fleet
Ceiling-raising code-quality elevation sweep across the standing codebase — compose the eleven -craft skills into ranked (scope, domain) targets, confirm one batch in a single up-front round that carries a taste-calibration sample of verbatim findings, fan out worktree-isolated subagents that each run the real harness-refactoring pipeline over one target's cited findings, admit nothing that lacks a cited craft finding and that a re-critique cannot show as net better, and hand back a tiered batch of elevation PRs and filed roadmap items for one bulk review. The fleet never auto-merges and never trusts a subagent's self-report.
The harness has a complete floor and no way to harvest its ceiling at batch scale. cleanup-fleet works the rule-based entropy queue — dead code, drift, structural risk in high-churn areas. bug-fleet hunts latent defects behind a reproduction bar. test-fleet chases coverage gaps. Every one of them acts on findings a machine can prove. Meanwhile the eleven -craft skills — naming-craft, code-craft, copy-craft, test-craft, spec-craft, docs-craft, knowledge-craft, api-craft, cli-ergonomics-craft, security-craft, harness-design-craft — encode the taste that says whether working code is any good, and they are invoked one file at a time, by a human who already suspected something was mediocre. The judgment exists; nothing sweeps with it.
craft-fleet is the ceiling twin of cleanup-fleet: it sweeps with the craft skills, ranks what they find, and hands back a tiered batch — bounded, high-confidence polish as elevation PRs, larger structural quality debt filed as roadmap items for the normal pipeline to build later. The restraint in that split is the design, not a shortfall of it. Craft findings are advisory LLM judgment by design, so a fleet that autonomously rewrites subjective "low quality" across a codebase produces churn, style-thrash, and bulk PRs that are miserable to review — it would spend the human's attention rather than save it, which inverts the entire point of the family. This member therefore leans file-don't-rewrite for anything structural, reserves direct PRs for safe, bounded, high-confidence polish, and strengthens the human-taste gate beyond every sibling's. It is a quality-queue member of the -fleet family: it does not sit on the core intake → decide → build → land spine, but works the craft-finding queue alongside it, exactly as cleanup-fleet works the entropy queue it mirrors.
This skill builds on the shared -fleet spine documented in docs/reference/fleet-family.md — the five-phase SELECT → CONFIRM → DISPATCH → VERIFY → terminal skeleton, the concurrency governor, the artifact + all-OS-CI verification discipline, the worktree fan-out with its nested-path push caveat, the front-load / park-unforeseen interaction model, and the never-silent-merge invariant. The family ADRs cited there — Subagent worktree fan-out (vs the Workflow primitive) for -fleet execution and The front-load / park-unforeseen interaction model for the -fleet family — state that contract once for the family, and the craft skills' shared 3-axis (tier × impact × confidence) output model supplies the finding vocabulary this skill consumes without extending. This SKILL.md defines only what is craft-fleet's own: its queue, its elevate-vs-file taxonomy, its cited-and-net-better verification, its tiered dual terminal act, and its domain-specific rationalizations.
When to Use
- Sweeping a codebase with the craft skills at batch scale, where per-file critique does not scale and the human's attention is the bottleneck
- Turning an existing craft-skill inventory into delivered elevation, rather than a standing list of things that could be better
- When the targets are genuinely independent — each is one coherent scope paired with one craft domain, elevated in its own worktree, and one target's polish does not depend on another's merge
- When the output must be trustworthy enough to review in bulk: every item arrives carrying the cited craft finding that produced it, so the reviewer judges taste rather than re-deriving it
- NOT for critiquing a single file — invoke the craft skill directly; a fleet's overhead only pays off across a batch
- NOT for rule-based entropy, dead code, or structural drift — that queue is
cleanup-fleet's, and it is objective where this one is advisory
- NOT for latent correctness defects — that is
bug-fleet; a craft finding that turns out to be a real bug is routed, never fixed here
- NOT for coverage gaps — closing them is
test-fleet; craft-fleet critiques the quality of the tests that exist, it does not author the ones that are missing
- NOT for landing or merging PRs — that is
pr-fleet; craft-fleet stops at reviewable and never merges
- NOT for applying a security fix — a
security-craft finding is never elevated by this fleet, and a genuine vulnerability is routed privately to the human rather than patched
- NOT for converging one target to clean — iterating a single module until it is good is a pipeline, not a fleet (which fans out across many independent targets into many outcomes)
Capability Roles
- Defines (Service Definition): the shared LLM-judgment-critique contract in
packages/cli/src/shared/craft/ — the LlmProvider interface (llm/provider.ts / llm/contracts.ts) plus the shared finding/axes schema (findings/axes.ts) and run store (runs/store.ts) — that every *-craft skill's critique phase conforms to. craft-fleet consumes this contract; it does not own it.
- Provides (Provider): the eleven craft skills —
naming-craft, spec-craft, code-craft, security-craft, test-craft, api-craft, cli-ergonomics-craft, copy-craft, docs-craft, harness-design-craft, knowledge-craft.
- Consumes (Consumer): this skill — craft-fleet composes ranked (scope, domain) targets across all craft providers uniformly through the shared critique/finding shape, then verifies each item by critique provenance. It is itself a
-fleet member and so is also a Provider to the fleet-command conductor.
Flags
| Flag |
Effect |
--concurrency |
Cap concurrent elevation subagents (default 2, max recommended 3 — the machine-storm limit) |
--domains |
Restrict the sweep to a comma-separated subset of the craft domains; the enabled set is confirmed at CONFIRM either way |
--report-only |
Compose the critique, rank the targets, and present the batch with its taste-calibration sample; do not dispatch elevations or file items |
--dry-run |
Run SELECT and CONFIRM only; stop before fan-out — nothing is verified, filed, or opened as a PR |
--file-only |
Open no elevation PR: every elevate target converts its budgeted elevation slot into a filed item carrying its cited finding, so the caps still hold |
Process
Iron Law
CITED-AND-NET-BETTER — no line is rewritten without a cited craft finding (a runId + rubricId + location from an actual craft-skill run), and nothing is emitted that a re-critique does not show as net better. The fleet never auto-applies a structural, contract-touching, or cross-module change, never elevates a prose file or a published contract, never publishes a routed security vulnerability, and never accepts a subagent's self-report as proof its pipeline ran.
The craft skills are advisory by design — they emit judgment carrying a visible confidence axis precisely because their findings are not binary — so a fleet built on them must not convert advice into authority. The cite is the ceiling analogue of a reproduction: it is the one piece of evidence that cannot be produced by asserting it. Confident prose explaining why a rewrite is better reads exactly like confident prose explaining why a rewrite that is worse is better; a runId and a rubricId either point at a real catalog run at a real location or they do not.
The second half of the law does the work the first cannot. A cite proves the location was worth looking at; it says nothing about whether the rewrite improved it. Only a re-critique of the changed code can distinguish elevation from style-thrash — swapping one finding for another is motion, not progress. And keeping the elevate boundary mechanical — high confidence, bounded, behavior-preserving, on an eligible surface — is what stops the whole thing degrading into "this rewrite felt safe," which is precisely how a taste-driven fleet becomes a churn engine.
The corollary matters as much as the law. A quiet target is a valid, valuable result. A target whose critique yields nothing above the noise floor tells the human where the ceiling has already been reached, and that is worth knowing. The pressure to manufacture an elevation so a sweep does not look wasted is the exact failure mode a subjective-judgment fleet must design against: there is no reproduction to fail here and no detector to stay red, so nothing but this rule stands between a thin batch and an invented one.
Phase 1: SELECT --> Phase 2: CONFIRM --> Phase 3: DISPATCH
|
v
Phase 5: FILE-AND-REPORT <-- Phase 4: VERIFY
| Phase |
Purpose |
Exit Condition |
| 1. SELECT |
Compose the craft skills into ranked (scope, domain) targets |
Ranked Target[] with routing verdicts, cross-check results, and floor/cap counts |
| 2. CONFIRM |
One human round: domains, batch, taste sample, floor, caps, governor, pinned base SHA |
Approved batch with a pinned base SHA and confirmed domains, floor, and caps |
| 3. DISPATCH |
Subagents run the real harness-refactoring over one target's cited findings |
Every elevate target returned a branch, downgraded, parked, or failed (all recorded) |
| 4. VERIFY |
Elevations: critique + elevation provenance, two-run re-critique, all-OS CI; filings: cite + cross-check |
Each item marked verified-elevation / verified-filing / routed / downgraded / rejected |
| 5. FILE-AND-REPORT |
Tiered dual terminal act, batch summary |
Report delivered; nothing merged |
Phase 1: SELECT — Compose the Craft Skills, Floor, Route, Rank
Compose the enabled craft skills — reimplement no critique. Run them over the repository. They already discover their own corpora and emit structured findings, each carrying a cite.rubricId, under a runId reported once per run in the run summary. A finding's cite is therefore composed — the run's runId paired with that finding's rubricId and location — rather than read off the finding alone. The fleet's job here is composition and ranking, not detection: it never re-derives a rubric, never invents a finding, and never restates a critique in its own words.
A missing or erroring craft skill degrades to the remaining ones and is recorded in the batch summary, never aborting the batch. If no craft skill is available, stop and report — there is nothing to rank.
Fold the findings into targets. A target is one coherent scope — a module, a doc set, a spec set — paired with exactly one craft domain. That pairing is not bookkeeping: it is what makes the terminal act's one-PR-per-target rule mean "never two craft domains in one PR," which is the property that keeps bulk taste review tractable. Two domains over the same scope are two targets.
Apply the noise floor. Drop, count, and never file or elevate any finding whose impact is small and which is additionally either tier: aspirational or confidence: low — that is, small ∧ (aspirational ∨ low). The grouping is stated explicitly because the rule must not be read as (small ∧ aspirational) ∨ low: a large-impact, low-confidence finding is exactly the kind of observation worth filing, and the wrong reading would silently discard it. Dropped findings are counted and reported, never silently discarded.
Cross-check every target. Search open elevation PRs and already-filed quality items for one that already addresses this target. An already-addressed target is dropped and annotated citing the resolving PR or item, never re-elevated. A fleet whose output is half duplicates costs more attention than it saves, by exactly the route its noise floor was meant to close.
Route every survivor mechanically — on the finding's own axes plus a surface rule, never on how bad the finding feels:
elevate (a direct PR) requires all of: confidence: high; the change is confined to one target and is behavior-preserving; no public-API, observable-contract, or exported-identifier change; no cross-module reach; and the finding's surface is on the elevation-eligible list for its domain — see Elevation Eligibility below.
file (a roadmap item) — everything else above the floor: any structural change, any contract-touching change, any medium- or low-confidence finding, and every finding in a file-only domain.
route — the finding is really a correctness defect or a genuine security vulnerability. Routing is park-and-hand-back, not a new mechanism: the item is neither elevated nor filed by this fleet, and it surfaces in the batch report for the human to place.
Route per finding, then re-form the targets by verdict. The routing rules read a finding, not a target, so a (scope, domain) pair whose findings split across verdicts never becomes a mixed target. It yields at most one elevate target and at most one file item for that same scope and domain, each carrying only its own findings; routed findings leave the pair entirely and are parked. That re-forming is what makes every downstream unit unambiguous — DISPATCH runs over one elevate target's findings plural, and VERIFY assigns exactly one verdict per emitted item. It also fixes the accounting: the caps count emitted items — elevation PRs and filed items — never findings.
Enforce the caps after ranking. Default 20 filed items and 20 elevation PRs per batch, hard — counted in emitted items, per the re-forming rule above. The cap keeps the highest tier × impact and drops the rest as over-cap, reported with its count — never silently. The caps bound SELECT-time intake, so a target that later downgrades from elevate to file converts an already-budgeted elevation slot rather than adding new intake; a batch therefore never hands back more than the two caps together allow, whatever happens downstream.
The cap, not the floor, is the real guard. Filing opens a tracking issue per item, so an uncapped sweep is a tracker flood no five-cell floor rule can prevent: the surviving medium-confidence middle of the distribution is large and legitimately routes to file. The floor removes the obvious tail cheaply; claiming it prevents backlog spam would be overselling a five-cell rule against a twenty-seven-cell distribution.
Score and order by tier × impact. Reuse harness-roadmap-pilot-style impact scoring so the ordering is principled and reproducible rather than a matter of which finding read most sharply. confidence is the routing axis and is deliberately not folded into the score. The 3-axis output model exists precisely because collapsing these axes destroys the information a reviewer needs to prioritize — a fleet that invented a second severity vocabulary would be re-collapsing them, and would drift from the catalogs it consumes.
Build the Target record for each survivor:
Target {
domain, // exactly one craft domain
id, // target slug
scope, // the files / docs / specs it covers
findings, // each: runId, rubricId, location, tier, impact, confidence
score, // composite tier x impact
verdict, // "elevate" | "file" — uniform, per the re-forming rule
crossCheck, // "novel" | "already-addressed" + resolving PR/item
forks, // detected decision forks to surface at CONFIRM (may be empty)
}
And the Batch record that scopes the whole run, settled once at CONFIRM:
Batch {
domains, // the enabled craft domains
baseSha, // the pinned base SHA every target works against
floor, // the confirmed noise floor
caps, // the confirmed per-batch caps { elevate, file }
governor, // the confirmed concurrency (default 2, max ~3)
targets, // the confirmed targets, each with its verdict
}
Phase 2: CONFIRM — The Single Up-Front Human Gate [checkpoint:human-verify]
Present the whole batch in one round. This is the only guaranteed human touchpoint before batch review — everything downstream runs autonomously. Present, together, in a single surface:
- The enabled craft domains — which of the eleven ran at all, so the human can switch a domain off before it costs a single elevation.
- The ranked targets, highest score first, each with its
elevate / file split and its score basis: the tier and impact that ranked it, and the routing call that split it.
- The taste-calibration sample — a handful of real, verbatim findings, elevation and file alike, drawn from the actual critique run rather than paraphrased or summarized.
- The noise floor with its drop count, and the per-batch caps with the over-cap count they shed. Both are re-tunable here, once.
- The proposed concurrency (default 2, capped at ~3).
- The pinned base SHA the whole batch works against. The SELECT critique run is pinned to it so VERIFY's branch re-critique is a like-for-like comparison rather than a moving target across a multi-hour batch, and so the green-baseline precondition the elevation pipeline requires is evaluated once for the batch instead of drifting per target.
Why the sample, and why here. Every sibling's CONFIRM presents a ranked batch; this one additionally presents verbatim findings, because taste does not generalize. Counts tell a human how much work is proposed; only a sample tells them whether this sweep's taste matches theirs, and that is the question on which the whole batch's value turns. It is also the cheapest possible place to discover a mismatch: disagreeing with the sample costs one conversation before fan-out, while discovering the same mismatch at review costs a batch of PRs plus all the machine time that produced them.
The human approves, trims, disables domains, or re-tunes the floor — once. Batch approval, domain selection, the floor, and the caps all settle in this same gate. Front-loading the genuinely-ambiguous calls is what keeps the autonomous stretch from producing work the human would have declined. A domain the human switches off runs nowhere downstream; a floor the human raises applies to the whole batch.
From here it is autonomous. After this gate the fleet does not pause per target. The only thing that re-surfaces before FILE-AND-REPORT is a target that hits a genuinely-unforeseen fork mid-flight, and that parks only that one target without blocking the batch. Under --dry-run the skill stops at the end of this phase; under --report-only it presents this surface and stops without dispatching, verifying, or filing.
Phase 3: DISPATCH — Worktree Fan-Out With a Concurrency Governor
One worktree-isolated subagent per confirmed elevate target. file targets require no fan-out at all — their critique is already complete and their terminal act is a filing, so dispatching them would spend machine time to produce nothing new.
Each subagent runs the real harness-refactoring pipeline over its one target's cited findings: tests green before and after every change, harness validate plus harness check-deps per step, blast radius computed up front, one small change per commit, and that skill's own revert-if-the-refactoring-introduced-no-improvement rule. The subagent does not hand-edit and does not short-cut the pipeline — the step-granular commit trail the pipeline necessarily leaves behind is exactly what VERIFY checks for.
The anti-churn discipline this fleet needs is therefore already law inside the skill it composes, rather than a policy layered on top of a free-hand editor. Composing that skill also inherits its precondition, which is broader than the suite alone: it refuses to run against a failing suite and requires a baseline harness validate and harness check-deps that both pass before the first step. Elevation therefore assumes a clean baseline at the pinned base SHA on all three.
The subagent runs no re-critique. That proof belongs to VERIFY and to VERIFY alone. The subagent's job ends at pushing a branch that carries its commit trail and its cited findings; anything it concluded about its own work is a claim, and a claim is not what the fleet's verdicts rest on.
Downgrade rules. A target whose elevation turns out to need a structural change — cross-module reach, a contract or exported-identifier change, a module split, an abstraction redesign — downgrades itself to file and reports, rather than applying it. A target whose baseline is not clean at the pinned base — a red suite, or a failing harness validate, or a failing harness check-deps — takes the same path, since those three checks together are the elevation pipeline's entire safety net. A downgrade is a normal outcome, not a failure: the critique remains valid and still reaches the human as a filed item; only the autonomous rewrite is withheld.
Cap concurrency at the confirmed governor (default 2, max ~3) and at the per-batch caps. This is the machine-storm limit: beyond roughly three concurrent elevation agents the compound load produces flaky failures indistinguishable from real ones — and in a fleet whose net-improvement proof is an already-noisy oracle, manufactured noise on top of it is uniquely corrosive. Never raise the cap to "go faster."
Record an "assumptions made" note per target — the ranking basis it worked from, the routing call it inherited, the elevation scope it actually took, and what it deliberately left un-elevated. Bulk taste review is only trustworthy when the reviewer can see what was assumed and what was consciously not touched.
Park the unforeseen. A target that hits a genuinely-unforeseen fork — the scope turns out to span two domains, a cited location no longer exists at the pinned base, the finding contradicts another the same run produced — parks that one target and reports it. The rest of the batch continues uninterrupted.
Push-path caveat. A worktree created under a nested agent-config path breaks the local pre-push documentation gate: it self-excludes and scans zero files. Subagents push via the GitHub API or from a non-nested throwaway worktree. Never --no-verify — bypassing the gate defeats the verification the fleet's guarantees rest on.
Elevation Eligibility — Only a Surface the Test Suite Guards
The whole safety envelope of the elevation pipeline is the test suite plus harness check-deps: those are what make "behavior-preserving" a checkable claim rather than an assertion. That yields one rule, not eleven judgment calls — a surface is elevation-eligible only if it lives inside source the test suite exercises. The single exception is narrower, not looser: test files are the suite rather than exercised by it, so they qualify only under the assertion-freeze rule stated below the table.
| Craft domain |
Elevation-eligible surface |
Otherwise |
naming-craft |
Non-exported local identifiers only |
file |
code-craft |
Within-unit simplification and control-flow honesty, signature unchanged |
file |
copy-craft |
Internal-facing prose only — code comments and internal log lines |
file |
test-craft |
Test names and test-body clarity, every assertion expression byte-identical |
file |
docs-craft |
Nothing — prose has no test suite to guard it |
file |
knowledge-craft |
Nothing — prose has no test suite to guard it |
file |
spec-craft |
Nothing — prose has no test suite to guard it; a ratified ADR is never edited at all |
file |
harness-design-craft |
Nothing — no craft-driven write path exists |
file |
api-craft |
Nothing — every surface it critiques is a published contract |
file |
cli-ergonomics-craft |
Nothing — every surface it critiques is a published contract |
file |
security-craft |
Nothing — never elevated |
file / route |
Four of eleven domains clear the bar. The other seven are file-only, and each for a stated reason rather than caution in general.
The prose domains are cut deliberately — docs-craft, knowledge-craft, and spec-craft are the tempting case, because prose looks like the safest thing in a repository to improve. It is the opposite. No skill in the toolset applies prose-quality edits under a safety envelope, so elevating prose would mean free-hand rewriting text with no mechanical check that it did not make things worse — which is precisely the churn this member exists to avoid, dressed as the easy win. A ratified ADR is additionally out of bounds on its own terms: it is a historical record of a decision, not a document to be improved, and editing one rewrites the past.
Published contracts are never elevated. api-craft and cli-ergonomics-craft critique published contracts by definition, so renaming a flag or an endpoint is a breaking change wearing a quality argument. security-craft is never auto-applied, because a wrong "improvement" to security posture is worse than the mediocrity it replaced, and posture is exactly the kind of judgment whose failure mode is silent. harness-design-craft has no reachable write path — its own polish phase emits before-and-after sketches and never modifies source, and the rule-based design-drift remediation path consumes drift findings that carry no craft runId or rubricId, so wiring it in would require inventing the finding translation the Iron Law exists to forbid.
copy-craft's narrowing establishes the general principle: routing follows the surface, not the skill that surfaced the finding. An error message and a CLI output string are the same bytes on the user's screen whether copy-craft or cli-ergonomics-craft found them, so they get the same treatment — filed, because user-facing output is an observable contract and no contract-touching change is ever elevated regardless of which domain raised it. What stays eligible is genuinely internal: code comments, which cannot alter behavior at all and are therefore the safest edit in the repository, and internal log lines, which are diagnostic output no consumer depends on. This shrinks the elevation surface; that is the direction this member is designed to err in.
test-craft's narrowing breaks a circularity. The elevation pipeline proves behavior preservation with the test suite, so elevating tests is circular unless the change provably cannot alter what the suite checks. The rule that breaks the circle is mechanical: every assertion expression must be byte-identical before and after, the passing-test count must be unchanged, and the set of passing test IDs may differ only by the renames the elevation itself applied — a rename changes a test's ID by construction, so freezing the ID set outright would forbid the very change this row exists to permit. Renaming a test or clarifying its arrange/act body qualifies. Sharpening an assertion does not: that changes what is asserted, which is a real improvement and a file, not an elevation.
Worker handoff — return the canonical FleetHandoffRecord. When a worker finishes its target it hands the orchestrator exactly one FleetHandoffRecord (from @harness-engineering/types) — the ONE bounded envelope every -fleet member emits, so fleet-command parses any fleet's worker output uniformly instead of special-casing an ad hoc per-worker report shape. The record carries status (done | parked | blocked | failed), fleet, item, a one-line summary, an evidence[] of verifiable pointers (branch, PR, artifact path, CI check — exactly the references VERIFY re-checks), next_steps[], and, for any non-done status, a blocker. The orchestrator validates it with validateFleetHandoffRecord; a malformed or unknown-keyed record is rejected, never silently misread. See the canonical handoff record in docs/reference/fleet-family.md.
Phase 4: VERIFY — Three Independent Proofs, Never Self-Report
Why three proofs and not one. No single artifact covers this fleet's two distinct risks. One risk is an agent applying its own taste — a change no critique asked for, which a clean commit trail and a green suite would both wave through. The other is an elevation that makes things worse — a change a real finding did ask for, applied so that it trades the cited problem for a new one, which a cite and a trail would in turn both wave through. Each proof closes what the others leave open, so VERIFY checks all three, independently, for every elevation — the file tier is verified here too, but to the narrower standard stated below, because it never produced a branch to check. Never accept a subagent's self-report: "cited it, elevated it, re-critiqued it clean" is a claim to be checked, not a result.
Critique provenance — the change was asked for. Every changed location maps to a cited finding from a real craft-skill run (runId + rubricId + location). A changed location that maps to no finding is the orchestrator's own taste and is rejected, however good it looks. That last clause is load-bearing: a taste-driven fleet's most attractive failure is the improvement nobody requested, and it is attractive precisely because it does look good.
Elevation provenance — the real pipeline ran. Confirm the step-granular harness-refactoring commit trail on the branch: one small change per commit, suite green throughout. Absent trail = the real pipeline did not run = rejected, however well the final diff reads. A hand-applied patch that happens to match a cited finding proves nothing about the safety envelope it skipped.
Net-improvement evidence — the change helped. Re-run the same craft skill over the changed scope on the branch and require both halves: the cited findings are resolved, and no new finding at equal-or-higher tier was introduced. Tier ordering is the craft catalogs' own — foundational outranks polish, which outranks aspirational — so "equal-or-higher" is evaluated mechanically rather than by feel. A re-critique that trades one finding for another is style-thrash, not elevation, and it does not ship.
The oracle is non-deterministic, and the protocol says so. A re-critique is an LLM call: two runs over identical code can disagree, so a single run can both falsely reject a good elevation and falsely accept a bad one. The answer is a two-run protocol biased conservative on both sides:
- A cited finding counts as resolved only if it is absent from both runs. Unanimity is required to credit an improvement.
- A new equal-or-higher-tier finding blocks if it appears in either run. One sighting of a regression is enough to stop shipping the rewrite.
- When the two runs disagree, the elevation is not proven — but the underlying critique is still valid, so the item downgrades to
file rather than being discarded. Nothing is lost; only the autonomous rewrite is withheld.
A re-critique that cannot run at all — no provider configured, a budget-exceeded prompt collection, an erroring skill — is a different case with a deliberately different outcome. It produces no proof, so by the same rule that rejects a missing commit trail the item is rejected and retried once, an unavailable provider usually being transient. The distinction is not fussiness: a re-critique that never ran leaves nothing to re-examine, while one that ran and split has already produced its reading — enough to justify filing the critique, never enough to justify shipping the rewrite.
Two runs rather than three is a deliberate cost call — craft skills bill per LLM call and this fleet runs eleven of them across a repository — and downgrade-not-discard is what makes the cheaper protocol safe: an inconclusive oracle costs the batch a filed item, never a bad merge.
VERIFY owns both runs; the DISPATCH subagent runs none. A subagent that re-critiques its own branch and reports the outcome leaves the orchestrator reading a claim rather than checking a proof — exactly what the family's never-self-report invariant forbids. Having both the subagent and VERIFY re-run would satisfy the invariant at four runs per item. Concentrating both runs here keeps the total at two and makes them independent, so the cost argument and the invariant are satisfied by the same choice instead of traded against each other.
Plus behavior preservation and CI. The passing-test count does not decrease; no public-API or observable-contract change is present; CI is green on all three operating systems plus the enforce and harness checks. Green on one OS is not green. These, like the three proofs above, are elevation checks.
The file tier is verified too — to its own, narrower standard. A file item never enters DISPATCH, so it has no branch, no commit trail, no re-critique, and no CI: four of the checks above have nothing to read. verified-filing therefore requires exactly two things — critique provenance, the cite resolving to a real craft-skill run (runId + rubricId + location) exactly as an elevation's must, and the cross-check, confirming at verification time that the target is still not already addressed by an open elevation PR or an existing filed item. An item failing either is rejected rather than filed. Those are the only two claims a filing actually makes, and a stale cite or a duplicate filing costs the human precisely the attention this fleet exists to save.
Assign exactly one verdict per item: verified-elevation, verified-filing, routed, downgraded, or rejected. A rejected item is retried once; still failing, it is reported as rejected with its reason and the batch continues. Degradation never reaches VERIFY. A missing craft skill degrades in SELECT and a structural discovery downgrades in DISPATCH, but a proof that this phase's standard for that item's tier demands and that cannot be produced is a rejection, not a graceful skip. The asymmetry is the point: the earlier phases can afford to lose coverage, and this one cannot afford to lose rigor.
Phase 5: FILE-AND-REPORT — Tiered Dual Terminal Act, Never Merge
One elevation PR per verified (target, craft domain) — never one per finding, never mixed across domains — each carrying its cited findings and its assumptions-made note. Never merged. The granularity is chosen, not inherited: one-PR-per-item is right when each item is a distinct defect, but forty naming fixes as forty PRs is a denial-of-service on review, and forty mixed fixes in one PR forces the reviewer to switch judgment modes line by line. Homogeneous batching gives the reviewer one kind of taste question at a time over one coherent scope, which is the only shape in which bulk taste review is actually tractable. Landing the batch is the human's call, optionally via pr-fleet.
File each verified file item as a roadmap item — including every target downgraded from elevate, whose critique remains valid even though its rewrite was withheld. Each filed item carries its cite, rubric, and location, so the eventual builder acts on the original judgment instead of re-deriving a critique that has already been paid for once. The orchestrator writes each filed item's "assumptions made" note here, from SELECT's routing basis: the ranking basis that ordered it, the routing call that split it, and why it was filed rather than elevated. A downgraded target arrives carrying the note its subagent already wrote in DISPATCH; a pure-file item never had a subagent, so this is the phase that produces its note.
Park-and-hand-back the routed findings. A correctness candidate is handed back as a seed candidate for the correctness queue — proving a defect requires a reproduction, and this member has no machinery for one. A genuine security vulnerability goes privately to the human and is never opened as a public item: filing a security finding publicly is disclosure, and this fleet has no disclosure machinery, no severity rating, and no embargo. Ordinary security-craft posture findings that are not vulnerabilities take the file path like any other file-only domain.
Emit a one-row-per-item batch summary for bulk review:
| Item |
Target |
Domain |
Verdict |
PR / Filed item |
Cite |
Assumptions made |
Alongside the table, report every non-item outcome with its count and its reason: dropped-by-noise-floor findings, over-cap findings shed after ranking, downgraded targets with their downgrade reason, cross-check drops each citing its resolving PR or item, routed findings, and quiet targets. Quiet targets are reported as quiet — a valid outcome, not a failure.
Under --file-only, every elevate target files instead. No branch is pushed and no PR is opened: each elevate-routed target **c
…(truncated)
1---2name: craft-fleet3description: Craft Fleet4---5# Craft Fleet67> Ceiling-raising code-quality elevation sweep across the standing codebase — compose the eleven `-craft` skills into ranked `(scope, domain)` targets, confirm one batch in a single up-front round that carries a **taste-calibration sample of verbatim findings**, fan out worktree-isolated subagents that each run the **real** `harness-refactoring` pipeline over one target's cited findings, admit nothing that lacks a cited craft finding and that a re-critique cannot show as **net better**, and hand back a **tiered** batch of elevation PRs and filed roadmap items for one bulk review. The fleet never auto-merges and never trusts a subagent's self-report.89The harness has a complete **floor** and no way to harvest its **ceiling** at batch scale. `cleanup-fleet` works the rule-based entropy queue — dead code, drift, structural risk in high-churn areas. `bug-fleet` hunts latent defects behind a reproduction bar. `test-fleet` chases coverage gaps. Every one of them acts on findings a machine can prove. Meanwhile the eleven `-craft` skills — `naming-craft`, `code-craft`, `copy-craft`, `test-craft`, `spec-craft`, `docs-craft`, `knowledge-craft`, `api-craft`, `cli-ergonomics-craft`, `security-craft`, `harness-design-craft` — encode the taste that says whether working code is any _good_, and they are invoked one file at a time, by a human who already suspected something was mediocre. The judgment exists; nothing sweeps with it.1011`craft-fleet` is the **ceiling twin of `cleanup-fleet`**: it sweeps with the craft skills, ranks what they find, and hands back a **tiered** batch — bounded, high-confidence polish as elevation PRs, larger structural quality debt filed as roadmap items for the normal pipeline to build later. The restraint in that split is the design, not a shortfall of it. Craft findings are advisory LLM judgment **by design**, so a fleet that autonomously rewrites subjective "low quality" across a codebase produces churn, style-thrash, and bulk PRs that are miserable to review — it would spend the human's attention rather than save it, which inverts the entire point of the family. This member therefore leans **file-don't-rewrite** for anything structural, reserves direct PRs for safe, bounded, high-confidence polish, and strengthens the human-taste gate beyond every sibling's. It is a **quality-queue** member of the `-fleet` family: it does not sit on the core intake → decide → build → land spine, but works the craft-finding queue alongside it, exactly as `cleanup-fleet` works the entropy queue it mirrors.1213This skill builds on the shared `-fleet` spine documented in `docs/reference/fleet-family.md` — the five-phase SELECT → CONFIRM → DISPATCH → VERIFY → terminal skeleton, the concurrency governor, the artifact + all-OS-CI verification discipline, the worktree fan-out with its nested-path push caveat, the front-load / park-unforeseen interaction model, and the never-silent-merge invariant. The family ADRs cited there — _Subagent worktree fan-out (vs the Workflow primitive) for `-fleet` execution_ and _The front-load / park-unforeseen interaction model for the `-fleet` family_ — state that contract once for the family, and the craft skills' shared _3-axis (tier × impact × confidence) output model_ supplies the finding vocabulary this skill consumes without extending. This SKILL.md defines only what is `craft-fleet`'s own: its queue, its elevate-vs-file taxonomy, its cited-and-net-better verification, its tiered dual terminal act, and its domain-specific rationalizations.1415## When to Use1617- Sweeping a codebase with the craft skills at batch scale, where per-file critique does not scale and the human's attention is the bottleneck18- Turning an existing craft-skill inventory into **delivered** elevation, rather than a standing list of things that could be better19- When the targets are genuinely independent — each is one coherent scope paired with one craft domain, elevated in its own worktree, and one target's polish does not depend on another's merge20- When the output must be trustworthy enough to review in bulk: every item arrives carrying the cited craft finding that produced it, so the reviewer judges taste rather than re-deriving it21- NOT for critiquing a single file — invoke the craft skill directly; a fleet's overhead only pays off across a batch22- NOT for rule-based entropy, dead code, or structural drift — that queue is `cleanup-fleet`'s, and it is objective where this one is advisory23- NOT for latent correctness defects — that is `bug-fleet`; a craft finding that turns out to be a real bug is **routed**, never fixed here24- NOT for coverage gaps — closing them is `test-fleet`; `craft-fleet` critiques the quality of the tests that exist, it does not author the ones that are missing25- NOT for landing or merging PRs — that is `pr-fleet`; `craft-fleet` stops at reviewable and never merges26- NOT for applying a security fix — a `security-craft` finding is never elevated by this fleet, and a genuine vulnerability is routed privately to the human rather than patched27- NOT for converging one target to clean — iterating a single module until it is good is a **pipeline**, not a fleet (which fans out across many independent targets into many outcomes)2829## Capability Roles3031<!-- Capability seam: this skill participates in a real extension point whose three roles are named and concrete. A seam with only one role filled is accidental single-implementation lock-in. See harness-skill-authoring Phase 1C. -->3233- **Defines (Service Definition):** the shared LLM-judgment-critique contract in `packages/cli/src/shared/craft/` — the `LlmProvider` interface (`llm/provider.ts` / `llm/contracts.ts`) plus the shared finding/axes schema (`findings/axes.ts`) and run store (`runs/store.ts`) — that every `*-craft` skill's critique phase conforms to. craft-fleet consumes this contract; it does not own it.34- **Provides (Provider):** the eleven craft skills — `naming-craft`, `spec-craft`, `code-craft`, `security-craft`, `test-craft`, `api-craft`, `cli-ergonomics-craft`, `copy-craft`, `docs-craft`, `harness-design-craft`, `knowledge-craft`.35- **Consumes (Consumer):** **this skill** — craft-fleet composes ranked (scope, domain) targets across all craft providers uniformly through the shared critique/finding shape, then verifies each item by critique provenance. It is itself a `-fleet` member and so is also a Provider to the `fleet-command` conductor.3637## Flags3839| Flag | Effect |40| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |41| `--concurrency` | Cap concurrent elevation subagents (default 2, max recommended 3 — the machine-storm limit) |42| `--domains` | Restrict the sweep to a comma-separated subset of the craft domains; the enabled set is confirmed at CONFIRM either way |43| `--report-only` | Compose the critique, rank the targets, and present the batch with its taste-calibration sample; do not dispatch elevations or file items |44| `--dry-run` | Run SELECT and CONFIRM only; stop before fan-out — nothing is verified, filed, or opened as a PR |45| `--file-only` | Open no elevation PR: every `elevate` target converts its budgeted elevation slot into a filed item carrying its cited finding, so the caps still hold |4647## Process4849### Iron Law5051**CITED-AND-NET-BETTER — no line is rewritten without a cited craft finding (a `runId` + `rubricId` + location from an actual craft-skill run), and nothing is emitted that a re-critique does not show as net better. The fleet never auto-applies a structural, contract-touching, or cross-module change, never elevates a prose file or a published contract, never publishes a routed security vulnerability, and never accepts a subagent's self-report as proof its pipeline ran.**5253The craft skills are advisory **by design** — they emit judgment carrying a visible confidence axis precisely because their findings are not binary — so a fleet built on them must not convert advice into authority. The cite is the ceiling analogue of a reproduction: it is the one piece of evidence that cannot be produced by asserting it. Confident prose explaining why a rewrite is better reads exactly like confident prose explaining why a rewrite that is worse is better; a `runId` and a `rubricId` either point at a real catalog run at a real location or they do not.5455The second half of the law does the work the first cannot. A cite proves the location was worth looking at; it says nothing about whether the rewrite improved it. Only a re-critique of the changed code can distinguish elevation from style-thrash — swapping one finding for another is motion, not progress. And keeping the elevate boundary **mechanical** — high confidence, bounded, behavior-preserving, on an eligible surface — is what stops the whole thing degrading into "this rewrite felt safe," which is precisely how a taste-driven fleet becomes a churn engine.5657The corollary matters as much as the law. **A quiet target is a valid, valuable result.** A target whose critique yields nothing above the noise floor tells the human where the ceiling has already been reached, and that is worth knowing. The pressure to manufacture an elevation so a sweep does not look wasted is the exact failure mode a subjective-judgment fleet must design against: there is no reproduction to fail here and no detector to stay red, so nothing but this rule stands between a thin batch and an invented one.5859```60Phase 1: SELECT --> Phase 2: CONFIRM --> Phase 3: DISPATCH61 |62 v63 Phase 5: FILE-AND-REPORT <-- Phase 4: VERIFY64```6566| Phase | Purpose | Exit Condition |67| ------------------ | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |68| 1. SELECT | Compose the craft skills into ranked `(scope, domain)` targets | Ranked `Target[]` with routing verdicts, cross-check results, and floor/cap counts |69| 2. CONFIRM | One human round: domains, batch, taste sample, floor, caps, governor, pinned base SHA | Approved batch with a pinned base SHA and confirmed domains, floor, and caps |70| 3. DISPATCH | Subagents run the real `harness-refactoring` over one target's cited findings | Every elevate target returned a branch, downgraded, parked, or failed (all recorded) |71| 4. VERIFY | Elevations: critique + elevation provenance, two-run re-critique, all-OS CI; filings: cite + cross-check | Each item marked verified-elevation / verified-filing / routed / downgraded / rejected |72| 5. FILE-AND-REPORT | Tiered dual terminal act, batch summary | Report delivered; nothing merged |7374### Phase 1: SELECT — Compose the Craft Skills, Floor, Route, Rank75761. **Compose the enabled craft skills — reimplement no critique.** Run them over the repository. They already discover their own corpora and emit structured findings, each carrying a `cite.rubricId`, under a `runId` reported once per run in the run summary. A finding's cite is therefore **composed** — the run's `runId` paired with that finding's `rubricId` and location — rather than read off the finding alone. The fleet's job here is composition and ranking, not detection: it never re-derives a rubric, never invents a finding, and never restates a critique in its own words.7778 A missing or erroring craft skill **degrades to the remaining ones and is recorded** in the batch summary, never aborting the batch. If no craft skill is available, stop and report — there is nothing to rank.79802. **Fold the findings into targets.** A **target** is one coherent scope — a module, a doc set, a spec set — paired with **exactly one craft domain**. That pairing is not bookkeeping: it is what makes the terminal act's one-PR-per-target rule mean "never two craft domains in one PR," which is the property that keeps bulk taste review tractable. Two domains over the same scope are two targets.81823. **Apply the noise floor.** Drop, count, and never file or elevate any finding whose `impact` is `small` **and** which is additionally either `tier: aspirational` **or** `confidence: low` — that is, `small` ∧ (`aspirational` ∨ `low`). The grouping is stated explicitly because the rule must not be read as (`small` ∧ `aspirational`) ∨ `low`: a `large`-impact, `low`-confidence finding is exactly the kind of observation worth filing, and the wrong reading would silently discard it. Dropped findings are **counted and reported**, never silently discarded.83844. **Cross-check every target.** Search open elevation PRs and already-filed quality items for one that already addresses this target. An already-addressed target is **dropped and annotated citing the resolving PR or item**, never re-elevated. A fleet whose output is half duplicates costs more attention than it saves, by exactly the route its noise floor was meant to close.85865. **Route every survivor mechanically** — on the finding's own axes plus a surface rule, never on how bad the finding feels:87 - **`elevate`** (a direct PR) requires **all** of: `confidence: high`; the change is confined to one target and is behavior-preserving; no public-API, observable-contract, or exported-identifier change; no cross-module reach; and the finding's surface is on the **elevation-eligible list** for its domain — see _Elevation Eligibility_ below.88 - **`file`** (a roadmap item) — everything else above the floor: any structural change, any contract-touching change, any `medium`- or `low`-confidence finding, and every finding in a file-only domain.89 - **`route`** — the finding is really a correctness defect or a genuine security vulnerability. Routing is **park-and-hand-back, not a new mechanism**: the item is neither elevated nor filed by this fleet, and it surfaces in the batch report for the human to place.9091 **Route per finding, then re-form the targets by verdict.** The routing rules read a finding, not a target, so a `(scope, domain)` pair whose findings split across verdicts never becomes a mixed target. It yields **at most one `elevate` target and at most one `file` item** for that same scope and domain, each carrying **only its own findings**; routed findings leave the pair entirely and are parked. That re-forming is what makes every downstream unit unambiguous — DISPATCH runs over one `elevate` target's findings plural, and VERIFY assigns exactly one verdict per emitted item. It also fixes the accounting: **the caps count emitted items — elevation PRs and filed items — never findings.**92936. **Enforce the caps after ranking.** Default **20 filed items and 20 elevation PRs** per batch, hard — counted in **emitted items**, per the re-forming rule above. The cap keeps the highest tier × impact and drops the rest **as over-cap, reported with its count** — never silently. The caps bound **SELECT-time intake**, so a target that later downgrades from `elevate` to `file` **converts an already-budgeted elevation slot** rather than adding new intake; a batch therefore never hands back more than the two caps together allow, whatever happens downstream.9495 **The cap, not the floor, is the real guard.** Filing opens a tracking issue per item, so an uncapped sweep is a tracker flood no five-cell floor rule can prevent: the surviving `medium`-confidence middle of the distribution is large and legitimately routes to `file`. The floor removes the obvious tail cheaply; claiming it prevents backlog spam would be overselling a five-cell rule against a twenty-seven-cell distribution.96977. **Score and order by tier × impact.** Reuse `harness-roadmap-pilot`-style impact scoring so the ordering is principled and reproducible rather than a matter of which finding read most sharply. `confidence` is the **routing** axis and is deliberately **not** folded into the score. The 3-axis output model exists precisely because collapsing these axes destroys the information a reviewer needs to prioritize — a fleet that invented a second severity vocabulary would be re-collapsing them, and would drift from the catalogs it consumes.98998. **Build the `Target` record** for each survivor:100101 ```102 Target {103 domain, // exactly one craft domain104 id, // target slug105 scope, // the files / docs / specs it covers106 findings, // each: runId, rubricId, location, tier, impact, confidence107 score, // composite tier x impact108 verdict, // "elevate" | "file" — uniform, per the re-forming rule109 crossCheck, // "novel" | "already-addressed" + resolving PR/item110 forks, // detected decision forks to surface at CONFIRM (may be empty)111 }112 ```113114 And the `Batch` record that scopes the whole run, settled once at CONFIRM:115116 ```117 Batch {118 domains, // the enabled craft domains119 baseSha, // the pinned base SHA every target works against120 floor, // the confirmed noise floor121 caps, // the confirmed per-batch caps { elevate, file }122 governor, // the confirmed concurrency (default 2, max ~3)123 targets, // the confirmed targets, each with its verdict124 }125 ```126127### Phase 2: CONFIRM — The Single Up-Front Human Gate `[checkpoint:human-verify]`1281291. **Present the whole batch in one round.** This is the **only guaranteed human touchpoint before batch review** — everything downstream runs autonomously. Present, together, in a single surface:130 - The **enabled craft domains** — which of the eleven ran at all, so the human can switch a domain off before it costs a single elevation.131 - The **ranked targets**, highest score first, each with its `elevate` / `file` split and its score basis: the tier and impact that ranked it, and the routing call that split it.132 - The **taste-calibration sample** — a handful of **real, verbatim findings**, elevation and file alike, drawn from the actual critique run rather than paraphrased or summarized.133 - The **noise floor** with its drop count, and the **per-batch caps** with the over-cap count they shed. Both are re-tunable here, once.134 - The **proposed concurrency** (default 2, capped at ~3).135 - The **pinned base SHA** the whole batch works against. The SELECT critique run is pinned to it so VERIFY's branch re-critique is a like-for-like comparison rather than a moving target across a multi-hour batch, and so the green-baseline precondition the elevation pipeline requires is evaluated once for the batch instead of drifting per target.1361372. **Why the sample, and why here.** Every sibling's CONFIRM presents a ranked batch; this one additionally presents verbatim findings, because **taste does not generalize**. Counts tell a human how much work is proposed; only a sample tells them whether this sweep's taste matches theirs, and that is the question on which the whole batch's value turns. It is also the cheapest possible place to discover a mismatch: disagreeing with the sample costs one conversation before fan-out, while discovering the same mismatch at review costs a batch of PRs plus all the machine time that produced them.1381393. **The human approves, trims, disables domains, or re-tunes the floor — once.** Batch approval, domain selection, the floor, and the caps all settle in this same gate. Front-loading the genuinely-ambiguous calls is what keeps the autonomous stretch from producing work the human would have declined. A domain the human switches off runs nowhere downstream; a floor the human raises applies to the whole batch.1401414. **From here it is autonomous.** After this gate the fleet does not pause per target. The only thing that re-surfaces before FILE-AND-REPORT is a target that hits a genuinely-unforeseen fork mid-flight, and that parks only that one target without blocking the batch. Under `--dry-run` the skill stops at the end of this phase; under `--report-only` it presents this surface and stops without dispatching, verifying, or filing.142143### Phase 3: DISPATCH — Worktree Fan-Out With a Concurrency Governor1441451. **One worktree-isolated subagent per confirmed `elevate` target.** `file` targets require no fan-out at all — their critique is already complete and their terminal act is a filing, so dispatching them would spend machine time to produce nothing new.1461472. **Each subagent runs the real `harness-refactoring` pipeline** over its one target's cited findings: tests green before and after **every** change, `harness validate` plus `harness check-deps` per step, blast radius computed up front, **one small change per commit**, and that skill's own revert-if-the-refactoring-introduced-no-improvement rule. The subagent does not hand-edit and does not short-cut the pipeline — the step-granular commit trail the pipeline necessarily leaves behind is exactly what VERIFY checks for.148149 The anti-churn discipline this fleet needs is therefore **already law inside the skill it composes**, rather than a policy layered on top of a free-hand editor. Composing that skill also inherits its precondition, which is broader than the suite alone: it refuses to run against a failing suite **and** requires a baseline `harness validate` and `harness check-deps` that both pass before the first step. Elevation therefore assumes a **clean baseline at the pinned base SHA** on all three.1501513. **The subagent runs no re-critique.** That proof belongs to VERIFY and to VERIFY alone. The subagent's job ends at pushing a branch that carries its commit trail and its cited findings; anything it concluded about its own work is a claim, and a claim is not what the fleet's verdicts rest on.1521534. **Downgrade rules.** A target whose elevation turns out to need a **structural** change — cross-module reach, a contract or exported-identifier change, a module split, an abstraction redesign — **downgrades itself to `file` and reports**, rather than applying it. A target whose **baseline is not clean** at the pinned base — a red suite, or a failing `harness validate`, or a failing `harness check-deps` — takes the same path, since those three checks together are the elevation pipeline's entire safety net. A downgrade is a **normal outcome, not a failure**: the critique remains valid and still reaches the human as a filed item; only the autonomous rewrite is withheld.1541555. **Cap concurrency at the confirmed governor (default 2, max ~3)** and at the per-batch caps. This is the machine-storm limit: beyond roughly three concurrent elevation agents the compound load produces flaky failures indistinguishable from real ones — and in a fleet whose net-improvement proof is an already-noisy oracle, manufactured noise on top of it is uniquely corrosive. Never raise the cap to "go faster."1561576. **Record an "assumptions made" note per target** — the ranking basis it worked from, the routing call it inherited, the elevation scope it actually took, and what it deliberately left un-elevated. Bulk taste review is only trustworthy when the reviewer can see what was assumed and what was consciously not touched.1581597. **Park the unforeseen.** A target that hits a genuinely-unforeseen fork — the scope turns out to span two domains, a cited location no longer exists at the pinned base, the finding contradicts another the same run produced — **parks that one target and reports it**. The rest of the batch continues uninterrupted.1601618. **Push-path caveat.** A worktree created under a nested agent-config path breaks the local pre-push documentation gate: it self-excludes and scans zero files. Subagents push via the GitHub API or from a non-nested throwaway worktree. **Never `--no-verify`** — bypassing the gate defeats the verification the fleet's guarantees rest on.162163### Elevation Eligibility — Only a Surface the Test Suite Guards164165The whole safety envelope of the elevation pipeline is the test suite plus `harness check-deps`: those are what make "behavior-preserving" a checkable claim rather than an assertion. That yields **one rule, not eleven judgment calls** — a surface is elevation-eligible only if it lives inside **source the test suite exercises**. The single exception is narrower, not looser: test files are the suite rather than exercised by it, so they qualify only under the assertion-freeze rule stated below the table.166167| Craft domain | Elevation-eligible surface | Otherwise |168| ---------------------- | ------------------------------------------------------------------------------------ | ------------ |169| `naming-craft` | Non-exported local identifiers only | file |170| `code-craft` | Within-unit simplification and control-flow honesty, signature unchanged | file |171| `copy-craft` | **Internal-facing prose only** — code comments and internal log lines | file |172| `test-craft` | Test names and test-body clarity, **every assertion expression byte-identical** | file |173| `docs-craft` | Nothing — prose has no test suite to guard it | file |174| `knowledge-craft` | Nothing — prose has no test suite to guard it | file |175| `spec-craft` | Nothing — prose has no test suite to guard it; a ratified ADR is never edited at all | file |176| `harness-design-craft` | Nothing — no craft-driven write path exists | file |177| `api-craft` | Nothing — every surface it critiques is a published contract | file |178| `cli-ergonomics-craft` | Nothing — every surface it critiques is a published contract | file |179| `security-craft` | Nothing — **never elevated** | file / route |180181Four of eleven domains clear the bar. The other seven are file-only, and each for a stated reason rather than caution in general.182183**The prose domains are cut deliberately** — `docs-craft`, `knowledge-craft`, and `spec-craft` are the tempting case, because prose looks like the safest thing in a repository to improve. It is the opposite. No skill in the toolset applies prose-quality edits under a safety envelope, so elevating prose would mean free-hand rewriting text with **no mechanical check that it did not make things worse** — which is precisely the churn this member exists to avoid, dressed as the easy win. A **ratified ADR is additionally out of bounds on its own terms**: it is a historical record of a decision, not a document to be improved, and editing one rewrites the past.184185**Published contracts are never elevated.** `api-craft` and `cli-ergonomics-craft` critique published contracts by definition, so renaming a flag or an endpoint is a breaking change wearing a quality argument. **`security-craft` is never auto-applied**, because a wrong "improvement" to security posture is worse than the mediocrity it replaced, and posture is exactly the kind of judgment whose failure mode is silent. **`harness-design-craft` has no reachable write path** — its own polish phase emits before-and-after sketches and never modifies source, and the rule-based design-drift remediation path consumes drift findings that carry no craft `runId` or `rubricId`, so wiring it in would require inventing the finding translation the Iron Law exists to forbid.186187**`copy-craft`'s narrowing establishes the general principle: routing follows the surface, not the skill that surfaced the finding.** An error message and a CLI output string are the same bytes on the user's screen whether `copy-craft` or `cli-ergonomics-craft` found them, so they get the same treatment — **filed**, because user-facing output is an observable contract and no contract-touching change is ever elevated regardless of which domain raised it. What stays eligible is genuinely internal: **code comments**, which cannot alter behavior at all and are therefore the safest edit in the repository, and **internal log lines**, which are diagnostic output no consumer depends on. This shrinks the elevation surface; that is the direction this member is designed to err in.188189**`test-craft`'s narrowing breaks a circularity.** The elevation pipeline proves behavior preservation **with** the test suite, so elevating tests is circular unless the change provably cannot alter what the suite checks. The rule that breaks the circle is mechanical: every assertion expression must be **byte-identical** before and after, the **passing-test count must be unchanged**, and the set of passing test IDs may differ **only by the renames the elevation itself applied** — a rename changes a test's ID by construction, so freezing the ID set outright would forbid the very change this row exists to permit. Renaming a test or clarifying its arrange/act body qualifies. **Sharpening an assertion does not**: that changes what is asserted, which is a real improvement and a `file`, not an elevation.190191**Worker handoff — return the canonical `FleetHandoffRecord`.** When a worker finishes its target it hands the orchestrator exactly one `FleetHandoffRecord` (from `@harness-engineering/types`) — the ONE bounded envelope every `-fleet` member emits, so `fleet-command` parses any fleet's worker output uniformly instead of special-casing an ad hoc per-worker report shape. The record carries `status` (`done | parked | blocked | failed`), `fleet`, `item`, a one-line `summary`, an `evidence[]` of verifiable pointers (branch, PR, artifact path, CI check — exactly the references VERIFY re-checks), `next_steps[]`, and, for any non-`done` status, a `blocker`. The orchestrator validates it with `validateFleetHandoffRecord`; a malformed or unknown-keyed record is rejected, never silently misread. See the canonical handoff record in `docs/reference/fleet-family.md`.192193### Phase 4: VERIFY — Three Independent Proofs, Never Self-Report1941951. **Why three proofs and not one.** No single artifact covers this fleet's two distinct risks. One risk is an agent applying **its own taste** — a change no critique asked for, which a clean commit trail and a green suite would both wave through. The other is an elevation that makes things **worse** — a change a real finding did ask for, applied so that it trades the cited problem for a new one, which a cite and a trail would in turn both wave through. Each proof closes what the others leave open, so VERIFY checks all three, independently, for **every elevation** — the `file` tier is verified here too, but to the narrower standard stated below, because it never produced a branch to check. **Never accept a subagent's self-report**: "cited it, elevated it, re-critiqued it clean" is a claim to be checked, not a result.1961972. **Critique provenance — the change was asked for.** Every changed location maps to a cited finding from a real craft-skill run (`runId` + `rubricId` + location). A changed location that maps to no finding is **the orchestrator's own taste and is rejected, however good it looks**. That last clause is load-bearing: a taste-driven fleet's most attractive failure is the improvement nobody requested, and it is attractive precisely because it does look good.1981993. **Elevation provenance — the real pipeline ran.** Confirm the step-granular `harness-refactoring` commit trail on the branch: one small change per commit, suite green throughout. **Absent trail = the real pipeline did not run = rejected**, however well the final diff reads. A hand-applied patch that happens to match a cited finding proves nothing about the safety envelope it skipped.2002014. **Net-improvement evidence — the change helped.** Re-run **the same craft skill** over the changed scope on the branch and require both halves: the cited findings are **resolved**, _and_ **no new finding at equal-or-higher tier was introduced**. Tier ordering is the craft catalogs' own — `foundational` outranks `polish`, which outranks `aspirational` — so "equal-or-higher" is evaluated mechanically rather than by feel. **A re-critique that trades one finding for another is style-thrash, not elevation**, and it does not ship.2022035. **The oracle is non-deterministic, and the protocol says so.** A re-critique is an LLM call: two runs over identical code can disagree, so a single run can both falsely reject a good elevation and falsely accept a bad one. The answer is a **two-run protocol biased conservative on both sides**:204 - A cited finding counts as **resolved only if it is absent from both runs**. Unanimity is required to credit an improvement.205 - A new equal-or-higher-tier finding **blocks if it appears in either run**. One sighting of a regression is enough to stop shipping the rewrite.206 - When the two runs **disagree**, the elevation is **not proven** — but the underlying critique is still valid, so the item **downgrades to `file`** rather than being discarded. Nothing is lost; only the autonomous rewrite is withheld.207208 A re-critique that **cannot run at all** — no provider configured, a budget-exceeded prompt collection, an erroring skill — is a different case with a deliberately different outcome. It produces **no proof**, so by the same rule that rejects a missing commit trail the item is **rejected and retried once**, an unavailable provider usually being transient. The distinction is not fussiness: a re-critique that never ran leaves nothing to re-examine, while one that ran and split has already produced its reading — enough to justify filing the critique, never enough to justify shipping the rewrite.209210 Two runs rather than three is a deliberate cost call — craft skills bill per LLM call and this fleet runs eleven of them across a repository — and downgrade-not-discard is what makes the cheaper protocol safe: **an inconclusive oracle costs the batch a filed item, never a bad merge.**2112126. **VERIFY owns both runs; the DISPATCH subagent runs none.** A subagent that re-critiques its own branch and reports the outcome leaves the orchestrator reading a **claim** rather than checking a **proof** — exactly what the family's never-self-report invariant forbids. Having both the subagent and VERIFY re-run would satisfy the invariant at four runs per item. Concentrating both runs here keeps the total at **two** _and_ makes them independent, so the cost argument and the invariant are satisfied by the same choice instead of traded against each other.2132147. **Plus behavior preservation and CI.** The passing-test count does not decrease; no public-API or observable-contract change is present; CI is green on **all three operating systems** plus the enforce and harness checks. Green on one OS is not green. These, like the three proofs above, are **elevation** checks.2152168. **The `file` tier is verified too — to its own, narrower standard.** A `file` item never enters DISPATCH, so it has no branch, no commit trail, no re-critique, and no CI: four of the checks above have nothing to read. `verified-filing` therefore requires exactly two things — **critique provenance**, the cite resolving to a real craft-skill run (`runId` + `rubricId` + location) exactly as an elevation's must, and the **cross-check**, confirming at verification time that the target is still not already addressed by an open elevation PR or an existing filed item. An item failing either is **rejected rather than filed**. Those are the only two claims a filing actually makes, and a stale cite or a duplicate filing costs the human precisely the attention this fleet exists to save.2172189. **Assign exactly one verdict per item:** `verified-elevation`, `verified-filing`, `routed`, `downgraded`, or `rejected`. A rejected item is **retried once**; still failing, it is reported as rejected with its reason and the batch continues. **Degradation never reaches VERIFY.** A missing craft skill degrades in SELECT and a structural discovery downgrades in DISPATCH, but a proof that this phase's standard for **that item's tier** demands and that cannot be produced is a **rejection**, not a graceful skip. The asymmetry is the point: the earlier phases can afford to lose coverage, and this one cannot afford to lose rigor.219220### Phase 5: FILE-AND-REPORT — Tiered Dual Terminal Act, Never Merge2212221. **One elevation PR per verified `(target, craft domain)`** — never one per finding, never mixed across domains — each carrying its cited findings and its assumptions-made note. **Never merged.** The granularity is chosen, not inherited: one-PR-per-item is right when each item is a distinct defect, but forty naming fixes as forty PRs is a denial-of-service on review, and forty mixed fixes in one PR forces the reviewer to switch judgment modes line by line. Homogeneous batching gives the reviewer **one kind of taste question at a time over one coherent scope**, which is the only shape in which bulk taste review is actually tractable. Landing the batch is the human's call, optionally via `pr-fleet`.2232242. **File each verified `file` item as a roadmap item** — including **every target downgraded from `elevate`**, whose critique remains valid even though its rewrite was withheld. Each filed item carries its **cite, rubric, and location**, so the eventual builder acts on the original judgment instead of re-deriving a critique that has already been paid for once. **The orchestrator writes each filed item's "assumptions made" note here**, from SELECT's routing basis: the ranking basis that ordered it, the routing call that split it, and why it was filed rather than elevated. A downgraded target arrives carrying the note its subagent already wrote in DISPATCH; a pure-`file` item never had a subagent, so this is the phase that produces its note.2252263. **Park-and-hand-back the routed findings.** A correctness candidate is handed back as a **seed candidate for the correctness queue** — proving a defect requires a reproduction, and this member has no machinery for one. A genuine security vulnerability goes **privately to the human** and is **never opened as a public item**: filing a security finding publicly _is_ disclosure, and this fleet has no disclosure machinery, no severity rating, and no embargo. Ordinary `security-craft` posture findings that are not vulnerabilities take the `file` path like any other file-only domain.2272284. **Emit a one-row-per-item batch summary** for bulk review:229230 | Item | Target | Domain | Verdict | PR / Filed item | Cite | Assumptions made |231 | ---- | ------ | ------ | ------- | --------------- | ---- | ---------------- |232233 Alongside the table, report every non-item outcome with its count and its reason: **dropped-by-noise-floor** findings, **over-cap** findings shed after ranking, **downgraded** targets with their downgrade reason, **cross-check drops** each citing its resolving PR or item, **routed** findings, and **quiet** targets. Quiet targets are reported **as quiet — a valid outcome, not a failure**.2342355. **Under `--file-only`, every `elevate` target files instead.** No branch is pushed and no PR is opened: each `elevate`-routed target **c236237…(truncated)