Agent concurrency
Claude Code's own ceilings are far higher than what is useful here: subagents
default to 20 concurrent (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS), nesting
runs 3 deep (CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH), the per-session spawn
cap was removed entirely in 2.1.224, and a workflow runs
min(16, CPUs − 2) agents at once. Nothing stops a fan-out from exhausting a
usage window in one turn.
These are the project's own limits. They are lower on purpose.
Caps
| Mode | Ceiling | Use when |
|---|---|---|
| Ultra / deep work | 5–6 agents | An audit, a multi-lens review, a wide research sweep. This is the maximum for any single fan-out. |
| Normal | 3–4 concurrent | Everything else, including background sessions running in parallel. |
Count concurrent, not total. Ten agents run three at a time is normal mode; six at once is ultra. If a plan needs more than six, it needs a second wave, not a bigger wave — and a second wave means the first one's results narrow the second, which is usually a better result anyway.
Never nest fan-outs. A depth of 3 means six agents each spawning six is
thirty-six live agents. Set CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH=1 if a task
tempts you.
Model and effort
Pick one of these three. Do not invent combinations.
| Tier | Use for |
|---|---|
| Opus 5 · xhigh | The hardest single agent in a fan-out — the judge, the synthesiser, the adversarial verifier. |
| Opus 5 · high | The default for everything else. |
| Opus 4.8 · xhigh | When a task benefits from 4.8's behaviour specifically, or Opus 5 is unavailable. |
Do not drop a fan-out to Haiku or Sonnet to "save budget" — a cheap agent that returns a wrong finding costs more than it saves, because the finding still has to be verified. Reduce the agent count instead.
Verify the model that actually ran before a model-specific claim or handoff.
model: frontmatter, --model, a saved preference, and a requested fallback
chain are intent, not execution evidence. [measured 2026-09-01] the
PreModelSwitch/PostModelSwitch hooks observed all six explicit interactive
switches but none of three successful unavailable-primary fallbacks. In an
interactive session read /status after the switch. In an unattended run use
the command or SDK result's actual model field. If neither readback exists, say
the model is unverified instead of naming the requested one as fact.
Effort is per-agent (effort: in skill frontmatter, opts.effort in a
workflow). Spend xhigh on the one agent whose judgement decides the outcome,
high on the rest.
Before spawning
- Say how many and why. "Six: one per audit dimension" is a plan. "Spawn agents to review this" is not.
- Check the budget. If the user set a
+Nktarget,budget.remaining()governs; a fan-out that would exhaust it should shrink, not proceed. - Prefer sequential when order matters. Parallel agents cannot see each other's findings; if agent 2's work depends on agent 1's, running them at once just produces two half-informed answers.
What an agent RETURNS is the cost, not which model ran it
[measured 2026-08-25] over 280 agents across 52 workflow runs: returns
totalled 3,524,077 characters, roughly 880k tokens fed back into main
threads. Median return 12,933 chars, p90 30,873, max 65,399. 60% exceed
10k. One run of 30 agents returned 658,588 chars — about 165k tokens — into a
single thread, which then re-reads them on every subsequent turn.
That is why model tier is the wrong lever. A main-thread request already re-reads ~405k tokens to emit ~1,000; the return is what grows that number permanently.
Write the artifact to a file; return the path and a summary. Agents that called Write or Edit returned a median 5,217 chars. Agents that wrote nothing returned 13,389 — 2.6× more. Only 20% of agents wrote anything, and the write-less 79% produced 86% of all returned characters.
The specific failure to avoid: 88 returns named a file path and still exceeded 10k chars. They wrote the file, then pasted the contents anyway. Naming the path is the point; the paste undoes it.
Budget each return explicitly in the prompt — 400–800 characters is enough for a path, a count, and the two or three things the caller must decide on. And note the 2,048-character cap is per schema string field, not a payload budget: 65,399-char returns exist, so a schema does not protect you. A field that truncates does so silently, mid-token, and the retry loop then burns five full generations against the same wall.
Never interpolate a large blob into a downstream prompt.
JSON.stringify(x, null, 1).slice(0, 90000) is 90k characters of prompt on
every call that touches it.
Serial chains duplicate; parallel ones do not
The intuition is backwards, and this corrects point 3 above rather than
replacing it. [measured] mean pairwise similarity between parallel agents'
returns was 0.008 (max 0.060 across 51 pairs). Serial chains averaged
0.072, with peaks of 0.710, 0.672 and 0.546 — and every pair above 0.25 sat
in a serial refine chain. One chain returned 113,915 characters re-emitting
substantially the same document fifteen times.
So sequencing is still right when order matters. But a serial stage must pass a delta, never the artifact — what changed and why, not the document again. The next stage can read the file.
Losing agents: it is the quota wall, not the width
[measured 2026-08-25] 42 of 280 agents (15%) were lost — 20,680 agent-seconds and 1,119
tool calls, journaled as nothing, because the journal records a result only on
completion. 20 of them carry a <synthetic> row reading "You've hit your
session limit", and 0 of those 20 journaled. That is 48% of all lost work from
one cause, and it is greppable after the fact.
Resist reading a loss rate by shape as evidence about width: 1-agent runs lost 67% while a 16-wide 30-agent run kept 30 of 30. Width sizes the blast radius of an interruption; it does not predict one. The caps above stand on usage limits and on second-wave-beats-bigger-wave, not on a loss rate.
The wall takes what is in flight, and the resume gives back what finished
[measured 2026-09-08] (docs/evidence-quota-wall-2026-09-08.md) the wall
lands on every agent in flight at once and on the main thread two seconds
later, so nothing can act at that moment. The recovery already exists:
Workflow({scriptPath, resumeFromRunId}) re-runs only the agent() calls whose
key has no result row in journal.jsonl and returns the rest from cache. On
the one real resume on this machine it re-started exactly the walled calls and
nothing else. Nothing calls it automatically, and the notification that names
it arrives while the thread is walled. Two things now close that gap:
scripts/workflow-run-triage.js <wf_id | --latest | --all-since <date>>prints, per agent, journaled / lost-quota-wall / lost-interrupted / lost-other with agent-seconds and tool calls, and the exact resume call.- The
stop-workflow-wall-note.jsStop hook says once, at the end of the first turn after the reset, that this session's latest run has walled agents, and puts the resume call in the model's context. It never blocks.
Resume; never relaunch. A fresh Workflow({script}) gets a new run id and
re-runs every agent. The session that made the one real resume wrote "nothing
cached, clean start" and was right only because nothing had finished.
Width is the bill for a wall, and serial is the bill for avoiding it
Both are measured, so pick with numbers rather than instinct. Over the 12 runs on this machine, wall-clock against the serial floor (the sum of agent-seconds):
| width in flight | runs | wall-clock as a fraction of serial |
|---|---|---|
| 3–4 | 8 | median 1 / 2.0 (serial costs 2.0×) |
| 8–12 | 4 | median 1 / 4.1 (serial costs 4.1×) |
And what a wall costs at a given width is width × the work each agent had done when it landed: the real wall took 5 agents at 40–59 s each, 190 agent-seconds in total, because it landed early; the same run's resume then lost one agent at 2,305 s to a user interrupt. So the rule is:
Anything that must not be lost runs in a serial chain, or in waves no wider than what you can afford to redo. A wide parallel phase is only for work that is cheap to re-run. With resume available, "afford to redo" means width × the per-agent cost, not the whole run; without it, it means the whole run.
The phase() convention that makes the choice visible before it runs: every
entry in meta.phases states its width and its re-run cost in detail, e.g.
{ title: 'Verify', detail: '2 wide, ~15 min each, must not be lost' } or
{ title: 'Scan', detail: '8 wide, ~1 min each, cheap to redo' }. A phase that
must not be lost and has more items than its width runs them in waves:
// Bounded blast radius: a wall takes at most `width` agents of this phase,
// and resumeFromRunId re-runs only those.
async function waves(items, width, fn) {
const out = []
for (let i = 0; i < items.length; i += width) {
const slice = items.slice(i, i + width)
out.push(...await parallel(slice.map((it, j) => () => fn(it, i + j))))
}
return out
}
Note that the built-in workflow-authoring reference (the Workflow tool's
script API) ships with Claude Code and is not in this repo, so this convention
lives here, where the **/*.workflow.js glob loads it.
The Task tools are gone on current models
As of 2.1.233, TaskCreate / TaskGet / TaskUpdate / TaskList and
TodoWrite are not available on Opus 4.8, Sonnet 5, Fable 5, and newer.
Any skill that lists them in allowed-tools will find them missing at run time.
Track multi-step work in prd.json — which is this framework's persistent task
system and survives /clear, compaction, and a restart, none of which the
native task list does. Set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 only if a user
explicitly wants the native list back.