Human Subagent: run the junior as an agent
The conversation's sharpest turn is an inversion of who imitates whom. Agents took the tactical seat: "tactical is the sergeant on the ground… strategic stuff is the general… agents are really good at tactical, really bad at strategic" (C25, quoted through the ledger). That seat was the entry-level job, and it was the on-ramp almost everyone used. You spent years being the sergeant, and somewhere in there you learned to be the general. Take the sergeant's work away and the on-ramp does not get shorter. It disappears, unless someone builds a replacement on purpose.
Bob's replacement is to hand the junior the agent's job description. At a company, "the lead engineer… should look at you as an agent and he should give you the same kind of tasks that the agents have and subject you to the same kind of deterministic tools… you should spend several months in that state being horribly unproductive but learning a hell of a lot. And by the time you've gone through that gauntlet, maybe you can be trusted to run an agent of your own" (C26). Matt Pocock names why the arithmetic works now: agents compress strategic feedback loops that used to take nine months, so a bad structural decision becomes visible fast enough to learn from (C26).
Naming. The word inside that quote is Bob's and is preserved as he said it. This pack does not adopt it: gauntlet is reserved for the Forge's gauntlet-loop, which means builder sub-agents fanned out against one falsifiable bar, shadowed by blind critics. That is agents producing artifacts. This is a person being drilled. Two different things, and one word for both would launder the difference. Here the practice is the drill, and the months are the rotation.
The ladder beneath it
The drill is the last rung, not the first. Bob's prerequisite is blunt: "you should be writing code for a year… so that you know what the agents are dealing with". Beneath that sits a descent-then-climb: "binary all the way through assembly language, some basic code like C, some higher level code like Python… and finally be able to strategically run an agent under supervision" (C26). The argument for the low rungs is not nostalgia but calibration: "if all you're doing is writing Java all day long, you live in a fantasy world" (C26).
Why the low rungs matter is this island's reading, not a claim C26 makes. A director who has never seen a pointer, a register, or a cache line has fewer independent ways to separate an agent's plausible answer from a correct one. The gates still refuse on their own authority, and that refusal needs no knowledge of registers. But everything the gates do not cover arrives as prose that sounds right, and reading that prose is where the calibration gets spent.
Read the ladder as entry criteria, not as a syllabus this island delivers. The rungs are Bob's (C26). The right-hand column is this island's reading of what each one buys, not a claim he made:
| Rung |
What it buys the future director |
| A year of writing code |
Knowing what the agents are dealing with (C26) |
| Binary, assembly |
Calibration: the machine stops being magic |
| A low-level language (C) |
Cost intuition: allocation, lifetime, failure |
| A high-level language (Python) |
Fluency in the layer the agents actually write |
| The drill |
Struggle under gates, supervised (C26) |
| Directing agents |
The strategic seat (C25) |
Where the rungs are already climbed, start at the drill. Where they are not, say so out loud rather than running a drill that measures the wrong gap.
Designing the drill
Six rules. They are what makes it a drill and not just a hard first quarter.
- Same brief, same shape. The trainee receives the mandate an agent would receive, in the same format: objective and definition of done, context pack, decision rights, stop conditions, evidence contract. That format is
delegated-authority-prompt's. Do not invent a gentler one for the human; the brief's gaps are half of what there is to learn.
- Same gates, unrelaxed, and say which lane you ran. The trainee finishes when the tool consents, not when a human says it looks fine. A gate is a loop nobody exits until the checker says okay (C4). Run the same gate stack the trainee will later direct (
crap-gate, mutant-hunt, dependency-fence). One honest tension: C17 makes a threshold a property of the executor, not of the value, so the agent lane sits looser than the human lane. That gap is about the agent's memory, not the trainee's comfort. Pick one lane, name it in the brief, and never retune it mid-rotation; a moving number teaches nothing.
- No rescue channel. No hint-dropping over the shoulder, no pair-programming the trainee out of the hole, no reviewer quietly fixing the diff. Every one of those is kindness that removes the struggle, and the struggle is the entire curriculum.
- Fresh brief per task, not a rolling epic. One task, one brief, one verdict: the born-do-die shape the agents run on (C10). Failure then lands on a task boundary where it can be examined, instead of smearing across a quarter.
- The rotation is time-boxed and disclosed. Bob's own words for the period are "horribly unproductive" (C26). Tell the trainee that before it starts, put an end date on it, and write both into the brief. Undisclosed, months of low output stop being a curriculum. They become a performance problem, and unfairly it becomes the trainee's.
- Book it outside the margin ledger.
margin-ledger exists to cut any gate that drags throughput below a human baseline (C5). A trainee under gates is designed to sit below that line. Log the rotation as training cost, explicitly, or the ledger will read a curriculum as a failing gate and correctly recommend cutting it.
Scoring: two boards, and only one is a number
Board 1, the gate tally (bookkeeping over enforced tools). Per task, record three things: which gate refused, how many fix-until-green cycles before it consented, and whether the trainee could state why it refused from the tool's output alone. The gates themselves are enforced by their own islands; this tally is not a gate and grants nothing. Its only job is to show the shape of the trainee's cycles over weeks. A falling cycle count on unchanged difficulty is progress; a flat one is a signal to look at the brief before blaming the trainee.
Board 2, the recognition board (the graduation criterion). Four behaviours, each scored observed or not yet, each stamped with the date and the task it was observed on:
- R1: Names a thrash signature in a running agent before being told, using the vocabulary of
thrash-watch.
- R2: Explains a refusal from the gate's own output, without a human translating it.
- R3: Chooses stop-and-repartition over one-more-fix at least once, and defends the choice afterward.
- R4: Re-derives a claimed verdict instead of accepting it. The first law, exercised on a colleague's assertion.
Two anti-laundering rules govern this board. First, not yet never softens into "observed weakly". Second, every mark traces to a real task the trainee ran, never to a hypothetical one. Where the behaviour reaches you as the trainee's own account — R2's explanation, R3's defence, Board 1's why — that account is checked against the artifacts of the task it describes, and it is corroboration rather than the observation itself. Graduation is observed on all four, which is to say the trainee has stopped being someone the gates catch and started being someone who reads the gates.
Why manufactured struggle is the point
The drill is not hazing and it is not a filter. It exists to produce one capability, and it is the capability a novice cannot fake. Bob knew his agents were failing in December because he had been there himself: "the important part was the next step where I watched them thrash. I could see the agent struggle and I recognized the struggle since I have been through that struggle… the novice would come in and not recognize the struggle" (C27).
Recognition is pattern-matching against your own scar tissue. C27 evidences that the novice arrives without it. The stronger claim, that no lecture installs it and no document transfers it (this one included), is this island's bet rather than his, and the whole drill is staked on that bet. If the tactical work that used to generate the scars is now done by agents, then the scars have to be manufactured deliberately, under supervision, on tasks small enough that the wreckage is legible.
This is also why the strategic seat cannot simply be assigned. Agents shrink accidental complexity; deciding what the system means stays human (seventies-canon, on Brooks's "No Silver Bullet"). The drill is how someone earns the right to that decision, and its exit criterion is Bob's, unchanged: trusted to run an agent of their own (C26).
Generalising past code
The drill transfers to any craft that has all three of these. Missing one, you have an ordinary apprenticeship. That is a good thing, but it is not this thing, and the scoring above will not fit it.
- A brief small enough to hand over whole. The unit of work fits in one mandate with a definition of done.
- A tool that can refuse. A check that says no on its own authority, so the loop closes without a human verdict (C4). Copy-edit linters, lab protocol validators, and CAD rule checks qualify outright: software says no, and the verdict does not depend on who ran it. A contract-review checklist is executed by a person, so it qualifies only when every hard-fail item is objectively decidable by any reader. The moment one item needs taste, you have "the senior thinks it's good" in a table. That is rule 3's rescue channel wearing a checklist, and it does not qualify.
- An expert who can already read the struggle. Someone who has been through it, per C27, and can score board 2 honestly.
Enforced vs advisory
advisory — everything this island prescribes. The six drill rules, both scoreboards, the graduation criterion, and the ladder's entry criteria are judgment work executed by a lead engineer. This island ships no script, deliberately: there is no honest mechanical check for "did this person learn to recognise struggle", and wrapping a thin gate around that judgment would be exactly the laundering the first law forbids. Anyone who tells you otherwise is selling a certificate.
enforced — the gates the trainee runs against are enforced by their own islands, with their own thresholds and their own red/green fixtures (known-dirty-fixture). The drill borrows that enforcement; it adds none.
enforced — this file's own shape, gated mechanically by the pack validator, which is the only claim on this page a machine checks.
python3 ../../scripts/validate-island.py ../human-subagent # exit 0
Run from this island's directory. It says this island is well-formed. It says nothing whatever about whether your drill works.
Boundaries: who owns what
- Lesson delivery is a Forge concern. Session state, lesson artifacts, learning records, the feedback loop that puts material in front of a person and remembers what they already did: all of it belongs to the Forge's
teach island, a stateful multi-session teaching workspace that is user-invoked — a person runs it. It is not one of the twenty-two boundaries COMPANION.md records, and it ships outside this repository, so what this pack can point you at is the ruling that drew the line, not the island: entry 33 of 03-FORGE50-AUDIT.md, which is why the drill design lives here and the delivery machinery does not. The concern is still not this island's. This island designs the drill and its scoring, and hands delivery there.
- The reading list is the sibling island. Which old books, and which load-bearing ideas each one carries for an agent-director, is
strategy-shelf's concern, and C27's cure for the novice. This island says the trainee needs the shelf; it never enumerates it.
- Recognising thrash live is
thrash-watch. That island owns the signatures and the intervention ladder. Board 2's R1 tests for that vocabulary; it does not redefine it.
- The ladder argument itself belongs to
abstraction-ladder: why the rung below is worth climbing, and the standing rung-below weekend for an incumbent director (C26, C28). This island reads the same C26 rungs one direction only, as one-time entry criteria for a trainee, scored on the way up.
- Not a
gauntlet. Restated because it matters: gauntlet-loop is agents fanned out against one bar with blind critics. This island is one human, one brief at a time, under the same tools. Never reuse the word for the drill.
- Whether to hire and train at all is
job-to-be-done's one-shot triage; this island starts after that answer is yes.
Done when
Nobody is handed the general's seat. The sergeant's work is gone, so the struggle has to be built on purpose. Recognising it is the one thing a novice cannot fake.
1---2name: human-subagent3description: Curriculum design for Uncle Bob's education inversion - a junior joining an agent-heavy team is run AS an agent, given the same task briefs and held to the same deterministic gates, spending months deliberately unproductive, until they can be trusted to direct agents of their own. Reach for it when onboarding a junior or a career-changer into an agent fleet, when planning what a new hire actually does for their first six months, or on "how do juniors learn anything now", "what do we do with new grads", "how should I train someone to run agents", "is there still an entry-level path". Differentiator - it designs the drill and its scoring only; the reading list is strategy-shelf, live struggle-spotting is thrash-watch, and the machinery that delivers lessons and remembers them belongs to the Forge.4---56# Human Subagent: run the junior as an agent78The conversation's sharpest turn is an inversion of who imitates whom. Agents took the tactical seat: *"tactical is the sergeant on the ground… strategic stuff is the general… agents are really good at tactical, really bad at strategic"* (C25, quoted through [the ledger](../../docs/01-CONCEPT-LEDGER.md)). That seat was the entry-level job, and it was the on-ramp almost everyone used. You spent years being the sergeant, and somewhere in there you learned to be the general. Take the sergeant's work away and the on-ramp does not get shorter. It disappears, unless someone builds a replacement on purpose.910Bob's replacement is to hand the junior the agent's job description. At a company, *"the lead engineer… should look at you as an agent and he should give you the same kind of tasks that the agents have and subject you to the same kind of deterministic tools… you should spend several months in that state being horribly unproductive but learning a hell of a lot. And by the time you've gone through that gauntlet, maybe you can be trusted to run an agent of your own"* (C26). Matt Pocock names why the arithmetic works now: agents compress strategic feedback loops that used to take nine months, so a bad structural decision becomes visible fast enough to learn from (C26).1112> **Naming.** The word inside that quote is Bob's and is preserved as he said it. This pack does not adopt it: `gauntlet` is reserved for the Forge's [`gauntlet-loop`](../../COMPANION.md#gauntlet-loop), which means builder sub-agents fanned out against one falsifiable bar, shadowed by blind critics. That is agents producing artifacts. This is a person being drilled. Two different things, and one word for both would launder the difference. Here the practice is **the drill**, and the months are **the rotation**.1314## The ladder beneath it1516The drill is the last rung, not the first. Bob's prerequisite is blunt: *"you should be writing code for a year… so that you know what the agents are dealing with"*. Beneath that sits a descent-then-climb: *"binary all the way through assembly language, some basic code like C, some higher level code like Python… and finally be able to strategically run an agent under supervision"* (C26). The argument for the low rungs is not nostalgia but calibration: *"if all you're doing is writing Java all day long, you live in a fantasy world"* (C26).1718Why the low rungs matter is this island's reading, not a claim C26 makes. A director who has never seen a pointer, a register, or a cache line has *fewer* independent ways to separate an agent's plausible answer from a correct one. The gates still refuse on their own authority, and that refusal needs no knowledge of registers. But everything the gates do not cover arrives as prose that sounds right, and reading that prose is where the calibration gets spent.1920Read the ladder as **entry criteria**, not as a syllabus this island delivers. The rungs are Bob's (C26). The right-hand column is this island's reading of what each one buys, not a claim he made:2122| Rung | What it buys the future director |23|---|---|24| A year of writing code | Knowing what the agents are dealing with (C26) |25| Binary, assembly | Calibration: the machine stops being magic |26| A low-level language (C) | Cost intuition: allocation, lifetime, failure |27| A high-level language (Python) | Fluency in the layer the agents actually write |28| The drill | Struggle under gates, supervised (C26) |29| Directing agents | The strategic seat (C25) |3031Where the rungs are already climbed, start at the drill. Where they are not, say so out loud rather than running a drill that measures the wrong gap.3233## Designing the drill3435Six rules. They are what makes it a drill and not just a hard first quarter.36371. **Same brief, same shape.** The trainee receives the mandate an agent would receive, in the same format: objective and definition of done, context pack, decision rights, stop conditions, evidence contract. That format is [`delegated-authority-prompt`](../../COMPANION.md#delegated-authority-prompt)'s. Do not invent a gentler one for the human; the brief's gaps are half of what there is to learn.382. **Same gates, unrelaxed, and say which lane you ran.** The trainee finishes when the tool consents, not when a human says it looks fine. A gate is a loop nobody exits until the checker says okay (C4). Run the *same* gate stack the trainee will later direct ([`crap-gate`](../crap-gate/SKILL.md), [`mutant-hunt`](../mutant-hunt/SKILL.md), [`dependency-fence`](../dependency-fence/SKILL.md)). One honest tension: C17 makes a threshold a property of the executor, not of the value, so the agent lane sits looser than the human lane. That gap is about the agent's memory, not the trainee's comfort. Pick one lane, name it in the brief, and never retune it mid-rotation; a moving number teaches nothing.393. **No rescue channel.** No hint-dropping over the shoulder, no pair-programming the trainee out of the hole, no reviewer quietly fixing the diff. Every one of those is kindness that removes the struggle, and the struggle is the entire curriculum.404. **Fresh brief per task, not a rolling epic.** One task, one brief, one verdict: the born-do-die shape the agents run on (C10). Failure then lands on a task boundary where it can be examined, instead of smearing across a quarter.415. **The rotation is time-boxed and disclosed.** Bob's own words for the period are *"horribly unproductive"* (C26). Tell the trainee that before it starts, put an end date on it, and write both into the brief. Undisclosed, months of low output stop being a curriculum. They become a performance problem, and unfairly it becomes the trainee's.426. **Book it outside the margin ledger.** [`margin-ledger`](../margin-ledger/SKILL.md) exists to cut any gate that drags throughput below a human baseline (C5). A trainee under gates is *designed* to sit below that line. Log the rotation as training cost, explicitly, or the ledger will read a curriculum as a failing gate and correctly recommend cutting it.4344## Scoring: two boards, and only one is a number4546**Board 1, the gate tally (bookkeeping over enforced tools).** Per task, record three things: which gate refused, how many fix-until-green cycles before it consented, and whether the trainee could state *why* it refused from the tool's output alone. The gates themselves are enforced by their own islands; this tally is not a gate and grants nothing. Its only job is to show the shape of the trainee's cycles over weeks. A falling cycle count on unchanged difficulty is progress; a flat one is a signal to look at the brief before blaming the trainee.4748**Board 2, the recognition board (the graduation criterion).** Four behaviours, each scored `observed` or `not yet`, each stamped with the date and the task it was observed on:4950- **R1**: Names a thrash signature in a running agent *before* being told, using the vocabulary of [`thrash-watch`](../thrash-watch/SKILL.md).51- **R2**: Explains a refusal from the gate's own output, without a human translating it.52- **R3**: Chooses stop-and-repartition over one-more-fix at least once, and defends the choice afterward.53- **R4**: Re-derives a claimed verdict instead of accepting it. The first law, exercised on a colleague's assertion.5455Two anti-laundering rules govern this board. First, `not yet` never softens into "observed weakly". Second, every mark traces to a real task the trainee ran, never to a hypothetical one. Where the behaviour reaches you as the trainee's own account — R2's explanation, R3's defence, Board 1's *why* — that account is checked against the artifacts of the task it describes, and it is corroboration rather than the observation itself. Graduation is `observed` on all four, which is to say the trainee has stopped being someone the gates catch and started being someone who reads the gates.5657## Why manufactured struggle is the point5859The drill is not hazing and it is not a filter. It exists to produce one capability, and it is the capability a novice cannot fake. Bob knew his agents were failing in December because he had been there himself: *"the important part was the next step where I watched them thrash. I could see the agent struggle and I recognized the struggle since I have been through that struggle… the novice would come in and not recognize the struggle"* (C27).6061Recognition is pattern-matching against your own scar tissue. C27 evidences that the novice arrives without it. The stronger claim, that no lecture installs it and no document transfers it (this one included), is this island's bet rather than his, and the whole drill is staked on that bet. If the tactical work that used to generate the scars is now done by agents, then the scars have to be manufactured deliberately, under supervision, on tasks small enough that the wreckage is legible.6263This is also why the strategic seat cannot simply be assigned. Agents shrink accidental complexity; deciding what the system *means* stays human ([seventies-canon](../../research/seventies-canon.md), on Brooks's "No Silver Bullet"). The drill is how someone earns the right to that decision, and its exit criterion is Bob's, unchanged: trusted to run an agent of their own (C26).6465## Generalising past code6667The drill transfers to any craft that has all three of these. Missing one, you have an ordinary apprenticeship. That is a good thing, but it is not this thing, and the scoring above will not fit it.68691. **A brief small enough to hand over whole.** The unit of work fits in one mandate with a definition of done.702. **A tool that can refuse.** A check that says no on its own authority, so the loop closes without a human verdict (C4). Copy-edit linters, lab protocol validators, and CAD rule checks qualify outright: software says no, and the verdict does not depend on who ran it. A contract-review checklist is executed by a person, so it qualifies only when every hard-fail item is objectively decidable by any reader. The moment one item needs taste, you have "the senior thinks it's good" in a table. That is rule 3's rescue channel wearing a checklist, and it does not qualify.713. **An expert who can already read the struggle.** Someone who has been through it, per C27, and can score board 2 honestly.7273## Enforced vs advisory7475- `advisory` — **everything this island prescribes.** The six drill rules, both scoreboards, the graduation criterion, and the ladder's entry criteria are judgment work executed by a lead engineer. This island ships *no script*, deliberately: there is no honest mechanical check for "did this person learn to recognise struggle", and wrapping a thin gate around that judgment would be exactly the laundering the first law forbids. Anyone who tells you otherwise is selling a certificate.76- `enforced` — the gates the trainee runs against are enforced by their own islands, with their own thresholds and their own red/green fixtures ([`known-dirty-fixture`](../known-dirty-fixture/SKILL.md)). The drill borrows that enforcement; it adds none.77- `enforced` — this file's own shape, gated mechanically by the pack validator, which is the only claim on this page a machine checks.7879```bash80python3 ../../scripts/validate-island.py ../human-subagent # exit 081```8283Run from this island's directory. It says this island is well-formed. It says nothing whatever about whether your drill works.8485## Boundaries: who owns what8687- **Lesson delivery is a Forge concern.** Session state, lesson artifacts, learning records, the feedback loop that puts material in front of a person and remembers what they already did: all of it belongs to the Forge's `teach` island, a stateful multi-session teaching workspace that is user-invoked — a person runs it. It is *not* one of the twenty-two boundaries [COMPANION.md](../../COMPANION.md) records, and it ships outside this repository, so what this pack can point you at is the ruling that drew the line, not the island: entry 33 of [03-FORGE50-AUDIT.md](../../docs/03-FORGE50-AUDIT.md), which is why the drill design lives here and the delivery machinery does not. The concern is still not this island's. This island designs the drill and its scoring, and hands delivery there.88- **The reading list is the sibling island.** Which old books, and which load-bearing ideas each one carries for an agent-director, is [`strategy-shelf`](../strategy-shelf/SKILL.md)'s concern, and C27's cure for the novice. This island says the trainee needs the shelf; it never enumerates it.89- **Recognising thrash live is [`thrash-watch`](../thrash-watch/SKILL.md).** That island owns the signatures and the intervention ladder. Board 2's R1 *tests* for that vocabulary; it does not redefine it.90- **The ladder argument itself** belongs to [`abstraction-ladder`](../abstraction-ladder/SKILL.md): why the rung below is worth climbing, and the standing rung-below weekend for an incumbent director (C26, C28). This island reads the same C26 rungs one direction only, as one-time *entry criteria* for a trainee, scored on the way up.91- **Not a `gauntlet`.** Restated because it matters: [`gauntlet-loop`](../../COMPANION.md#gauntlet-loop) is agents fanned out against one bar with blind critics. This island is one human, one brief at a time, under the same tools. Never reuse the word for the drill.92- **Whether to hire and train at all** is [`job-to-be-done`](../../COMPANION.md#job-to-be-done)'s one-shot triage; this island starts after that answer is yes.9394## Done when9596- [ ] The trainee's ladder position is stated, and any unclimbed rung is named rather than assumed.97- [ ] The brief uses the agents' mandate format, and the gate lane is named in it.98- [ ] The rotation has a disclosed end date and is booked as training cost outside the margin ledger.99- [ ] Board 1 has a per-task row; Board 2 has four rows, each `observed` with a task and a date, or `not yet`.100- [ ] Graduation was declared on four `observed` marks, never on a report of progress.101102**Nobody is handed the general's seat. The sergeant's work is gone, so the struggle has to be built on purpose. Recognising it is the one thing a novice cannot fake.**