Orchestration - Automated Multi-Agent Coordination
Scheduling
Goal
Automatically orchestrate multi-agent execution with task decomposition, native/fallback dispatch, memory coordination, progress monitoring, verification, QA cross-review, retry, and result collection.
Intent signature
- User asks to orchestrate, run in parallel, automate multi-agent execution, or coordinate full-stack work end to end.
- Task requires multiple specialist agents and a persistent review/remediation loop.
When to use
- Complex feature requires multiple specialized agents working in parallel
- User wants automated execution without manually spawning agents
- Full-stack implementation spanning backend, frontend, mobile, and QA
- User says "run it automatically", "run in parallel", or similar automation requests
When NOT to use
- Simple single-domain task -> use the specific agent directly
- User wants step-by-step manual control -> use oma-coordination
- Quick bug fixes or minor changes
Expected inputs
- Complex feature or workflow request
- Project config, model/vendor routing, agent types, task constraints, and workspace/session needs
- Acceptance criteria and verification expectations
Expected outputs
- Orchestrator session state, task board, progress files, result files, and final summary
- Specialist agent outputs after mechanical checks, automated verify, and QA cross-review
- Review history and retry/remediation status when loops fail
Dependencies
.agents/oma-config.yaml, .codex/agents/*.toml, .gemini/agents/*.md, or fallback oma agent:spawn
- Memory provider config, subagent prompt template, scripts, task templates, verify script, and session metrics
Control-flow features
- Branches by vendor/native dispatch availability, priority tiers, agent completion/failure, verification status, QA verdict, retry limits, and clarification debt
- Spawns processes/agents and reads/writes memory/result files
- Blocks termination until persistent workflows complete
Structural Flow
Entry
- Resolve agent vendor routing and runtime dispatch path.
- Decompose request into priority-tiered tasks.
- For each task, classify into one or more
domain_tags by matching against the Intent signature block of each installed .agents/skills/oma-*/SKILL.md. Tasks that match no domain confidently inherit the union of their parent feature's tags.
- Build a per-task
exposed_skill_set = skills whose name is in domain_tags. If |exposed_skill_set| < 2 after classification, fall back to the full installed set (flat exposure) and record exposure_fallback: true in the task board.
- Create session memory and task board with
exposed_skill_set and exposure_fallback per task.
Scenes
- PREPARE: Plan, setup session ID, and initialize memory files.
- ACT: Spawn agents by priority tier within parallelism limits.
- VERIFY: Run self-check,
oma verify, and QA cross-review loop.
- RECOVER: Retry failed agents with review history when limits allow.
- FINALIZE: Collect result files, compile summary, and clean progress files.
Transitions
- If native dispatch is available for current runtime/vendor, use it.
- If vendors differ or native path is unavailable, use fallback spawn.
- If verify or QA fails, feed feedback back to the implementation agent.
- If review loop limits are exceeded, report review history and quality warning.
- If a task's
exposed_skill_set excludes a skill that a recovered failure indicates was needed, re-classify the task and re-dispatch with the expanded set rather than retrying against the original narrow set.
Failure and recovery
- Retry failed agents up to configured limits.
- Re-spawn with review history when review loop is exhausted.
- Pause or request re-specification when clarification debt thresholds are exceeded.
Exit
- Success: all tasks complete, verify/review pass, and results are summarized.
- Partial success: failed agents, exhausted review loops, or clarification debt are explicit.
Logical Operations
Actions
| Action |
SSL primitive |
Evidence |
| Read config and task context |
READ |
oma config, routing, request |
| Classify task into domain tags |
INFER |
task text vs each skill's Intent signature |
| Compute exposed skill set |
SELECT |
intersection of domain tags and installed skills |
| Select dispatch path |
SELECT |
Native vs fallback |
| Write session state |
WRITE |
task board and memory files |
| Spawn agents |
CALL_TOOL |
native CLI or oma agent:spawn |
| Poll progress |
READ |
progress/result files |
| Run verification |
CALL_TOOL |
oma verify, tests, QA |
| Update retry state |
UPDATE_STATE |
loop counters and CD metrics |
| Report final result |
NOTIFY |
compiled summary |
Tools and instruments
- Native CLI subagent dispatch, fallback spawn scripts, memory tools, verify script, QA agent
- Session metrics, prompt templates, task templates
Canonical command path
oma agent:spawn <agent-type> "<task>" <session-id> -w <workspace>
oma verify <agent-type> --workspace <workspace> --json
When native runtime dispatch is available, prefer the runtime-specific native path listed in this skill before falling back to oma agent:spawn.
Resource scope
| Scope |
Resource target |
LOCAL_FS |
Session, task-board, progress, result, config files |
PROCESS |
Agent CLI processes and verify scripts |
MEMORY |
Session state and clarification debt |
CODEBASE |
Workspaces owned by spawned agents |
Preconditions
- Task is decomposable into specialist agent work.
- Runtime/vendor dispatch path or fallback exists.
Effects and side effects
- Spawns agents and writes session/progress/result artifacts.
- May cause code changes through specialist agents.
- May trigger iterative review and retries.
Guardrails
- Orchestrate per-agent dispatch from the project configuration before spawning any agent.
- If
target_vendor === current_runtime_vendor and the runtime has a verified native path, use native dispatch.
- Otherwise fall back to
oma agent:spawn.
- Never exceed the configured parallelism or retry limits.
- Keep session state, task-board state, progress files, and result files aligned throughout the run.
- Domain gating must be soft: prefer a narrower
exposed_skill_set, but fall back to flat exposure when classification confidence is low rather than starving a task of a required specialist.
Current native executor paths:
- Claude Code: Agent tool with
.claude/agents/{agent}.md definitions (multiple Agent tool calls in one message run in parallel; results return synchronously — no polling)
- OpenCode: native
task tool with subagent_type: {agent-id}; do not use oma agent:spawn for same-session OpenCode work because it will not appear as a native child task
- Codex CLI:
codex exec "@agent ..." using .codex/agents/*.toml
- Gemini CLI:
gemini -p "@agent ..." using .gemini/agents/*.md
Vendor-specific execution protocols are injected automatically for fallback CLI runs.
Configuration
| Setting |
Default |
Description |
| MAX_PARALLEL |
3 |
Max concurrent subagents |
| MAX_RETRIES |
2 |
Retry attempts per failed task |
| POLL_INTERVAL |
30s |
Status check interval |
| MAX_TURNS (impl) |
20 |
Turn limit for backend/frontend/mobile |
| MAX_TURNS (review) |
15 |
Turn limit for qa/debug |
| MAX_TURNS (plan) |
10 |
Turn limit for pm |
These are skill-level defaults applied by the orchestrating agent; they are not read from config/cli-config.yaml (which carries only vendor CLI and execution settings such as results_dir and timeout).
Memory Configuration
Memory provider and tool names are configurable via .agents/mcp.json (not the repo-root .mcp.json, which is the Claude Code MCP server config):
{
"memoryConfig": {
"provider": "file",
"basePath": ".agents/state/memories",
"tools": {
"read": "Read",
"write": "Write",
"edit": "Edit"
}
}
}
Workflow Phases
PHASE 1 - Plan: Analyze request -> decompose tasks -> generate session ID
PHASE 1.5 - Domain gate: For each task, intersect Intent signature matches across installed skills to derive exposed_skill_set. Record exposure_fallback: true when the intersection is too small to be useful and the flat library is used instead.
PHASE 2 - Setup: Use memory write tool to create orchestrator-session.md + task-board.md (include exposed_skill_set per task)
PHASE 3 - Execute: Spawn agents by priority tier (never exceed MAX_PARALLEL); inject only exposed_skill_set into each subagent's available specialist list
PHASE 4 - Monitor: Poll every POLL_INTERVAL; handle completed/failed/crashed agents
PHASE 4.5 - Verify: Run mechanical checks for every completed agent; run oma verify {agent-type} only for backend, frontend, mobile, qa, debug, and pm; then run QA cross-review for every completed implementation
PHASE 5 - Collect: Read all result-{agent}-{sessionId}.md, compile summary, cleanup progress files
See resources/subagent-prompt-template.md for prompt construction.
See resources/memory-schema.md for memory file formats.
Memory File Ownership
| File |
Owner |
Others |
orchestrator-session.md |
orchestrator |
read-only |
task-board.md |
orchestrator |
read-only |
progress-{agent}[-{sessionId}].md |
that agent |
orchestrator reads |
result-{agent}[-{sessionId}].md |
that agent |
orchestrator reads |
Agent-to-Agent Review Loop (PHASE 4.5)
After each agent completes, enter an iterative review loop, not a single-pass verification.
Loop Flow
Agent completes work
↓
[1] Mechanical Self-Check: lint, type-check, tests, diff scope
↓
[2] Verify: For supported types, run `oma verify {agent-type} --workspace {workspace}`
Unsupported (`db`, `refactor`, `architecture`, `tf-infra`, `docs`) → record SKIP and continue
↓ FAIL → Agent receives feedback, fixes, back to [1]
↓ PASS
[3] Cross-Review: QA agent reviews the changes
↓ FAIL → Agent receives review feedback, fixes, back to [1]
↓ PASS
Accept result
Step Details
[1] Mechanical Self-Check (formerly "Self-Review"):
Before requesting external review, the implementation agent must:
- Run lint, type-check, and tests in the workspace
- Verify only planned files were modified (diff scope check)
- Fix any mechanical failures (compile errors, test failures)
Quality judgment is NOT performed in this step.
Design quality, architecture alignment, and acceptance criteria satisfaction
are evaluated exclusively in [3] Cross-Review by the QA agent.
Reason: Self-evaluation bias causes agents to consistently overrate their own output
(ref: Anthropic harness design research).
[2] Automated Verify:
oma verify {agent-type} --workspace {workspace} --json
- Run only for
backend, frontend, mobile, qa, debug, and pm.
- For
db, refactor, architecture, tf-infra, and docs, record that automated verify is unsupported and continue to QA cross-review after the mechanical checks.
- PASS (exit 0): Proceed to cross-review
- FAIL (exit 1): Feed verify output back to the agent as correction context
[3] Cross-Review: Spawn QA agent to review the changes:
- QA agent reads the diff, runs checks, evaluates against acceptance criteria
- If
docs/CODE-REVIEW.md exists, QA agent uses it as the review checklist
- QA agent outputs: PASS (with optional nits) or FAIL (with specific issues)
- On FAIL: issues are fed back to the implementation agent for fixing
Loop Limits
| Counter |
Max |
On Exceeded |
| Self-check + fix cycles |
3 |
Escalate to cross-review regardless |
| Cross-review rejections |
2 |
Report to user with review history |
| Total loop iterations |
5 |
Force-complete with quality warning |
Review Feedback Format
When feeding review results back to the implementation agent:
## Review Feedback (iteration {n}/{max})
**Reviewer**: {self / verify / qa-agent}
**Verdict**: FAIL
**Issues**:
1. {specific issue with file and line reference}
2. {specific issue}
**Fix instruction**: {what to change}
This replaces single-pass verification. Most "nitpicking" should happen agent-to-agent.
Human review is reserved for final approval, not catching lint errors.
Retry Logic (after review loop exhaustion)
Before starting any retry, check the termination conditions (OR, whichever fires first wins):
- Retry cap: retry count for this agent has reached MAX_RETRIES — do not start another cycle.
- Session cost cap: if a quota cap is configured (
loadQuotaCap() from cli/io/session-cost.ts; no cap → skip), call checkCap(sessionId, cap). On exceeded === true, save the agent's partial results, report early termination due to quota, and do not spawn the next retry or any remaining agents in the tier.
If neither condition fires:
- 1st retry: Re-spawn agent with full review history as context
- 2nd retry: Re-spawn with "Try a different approach" + review history
- After MAX_RETRIES exhausted (cost cap not exceeded): activate the Exploration Loop (see
orchestrate.md Step 5): generate 2-3 alternative hypotheses, spawn the same agent type with different hypothesis prompts in parallel separate workspaces, score with Quality Score when available, keep the highest-scoring approach, and record all experiments in the Experiment Ledger.
- Final failure: Report to user with complete review trail, ask whether to continue or abort
Clarification Debt (CD) Monitoring
Track user corrections during session execution. See ../_shared/core/session-metrics.md for full protocol.
Event Classification
When user sends feedback during session:
- clarify (+10): User answering agent's question
- correct (+25): User correcting agent's misunderstanding
- redo (+40): User rejecting work, requesting restart
Threshold Actions
| CD Score |
Action |
| CD >= 50 |
RCA Required: QA agent must add entry to lessons-learned.md |
| CD >= 80 |
Session Pause: Request user to re-specify requirements |
redo >= 2 |
Scope Lock: Request explicit allowlist confirmation before continuing |
Recording
After each user correction event:
[EDIT]("session-metrics.md", append event to Events table)
At session end, if CD >= 50:
- Include CD summary in final report
- Trigger QA agent RCA generation
- Update
lessons-learned.md with prevention measures
References
- Prompt template:
resources/subagent-prompt-template.md
- Memory schema:
resources/memory-schema.md
- Config:
config/cli-config.yaml
- Scripts:
scripts/spawn-agent.sh, scripts/parallel-run.sh, scripts/verify.sh
- Task templates:
templates/
- Skill-to-agent mapping:
../_shared/core/skill-routing.md
- Verification:
scripts/verify.sh <agent-type>
- Session metrics:
../_shared/core/session-metrics.md
- API contract template (SSOT):
../_shared/core/api-contracts/template.md; read generated contracts from .agents/results/api-contracts/ (run artifact) or docs/plans/contracts/ (durable spec)
- Context loading:
../_shared/core/context-loading.md
- Difficulty guide:
../_shared/core/difficulty-guide.md
- Clarification protocol:
../_shared/core/clarification-protocol.md
- Context budget:
../_shared/core/context-budget.md
- Lessons learned:
../_shared/core/lessons-learned.md
1---2name: oma-orchestration-33description: Automated multi-agent orchestration that spawns CLI subagents in parallel, coordinates via MCP Memory, and monitors progress. Use for orchestration, parallel execution, and automated multi-agent workflows.4---56# Orchestration - Automated Multi-Agent Coordination78## Scheduling910### Goal11Automatically orchestrate multi-agent execution with task decomposition, native/fallback dispatch, memory coordination, progress monitoring, verification, QA cross-review, retry, and result collection.1213### Intent signature14- User asks to orchestrate, run in parallel, automate multi-agent execution, or coordinate full-stack work end to end.15- Task requires multiple specialist agents and a persistent review/remediation loop.1617### When to use18- Complex feature requires multiple specialized agents working in parallel19- User wants automated execution without manually spawning agents20- Full-stack implementation spanning backend, frontend, mobile, and QA21- User says "run it automatically", "run in parallel", or similar automation requests2223### When NOT to use24- Simple single-domain task -> use the specific agent directly25- User wants step-by-step manual control -> use oma-coordination26- Quick bug fixes or minor changes2728### Expected inputs29- Complex feature or workflow request30- Project config, model/vendor routing, agent types, task constraints, and workspace/session needs31- Acceptance criteria and verification expectations3233### Expected outputs34- Orchestrator session state, task board, progress files, result files, and final summary35- Specialist agent outputs after mechanical checks, automated verify, and QA cross-review36- Review history and retry/remediation status when loops fail3738### Dependencies39- `.agents/oma-config.yaml`, `.codex/agents/*.toml`, `.gemini/agents/*.md`, or fallback `oma agent:spawn`40- Memory provider config, subagent prompt template, scripts, task templates, verify script, and session metrics4142### Control-flow features43- Branches by vendor/native dispatch availability, priority tiers, agent completion/failure, verification status, QA verdict, retry limits, and clarification debt44- Spawns processes/agents and reads/writes memory/result files45- Blocks termination until persistent workflows complete4647## Structural Flow4849### Entry501. Resolve agent vendor routing and runtime dispatch path.512. Decompose request into priority-tiered tasks.523. For each task, classify into one or more `domain_tags` by matching against the `Intent signature` block of each installed `.agents/skills/oma-*/SKILL.md`. Tasks that match no domain confidently inherit the union of their parent feature's tags.534. Build a per-task `exposed_skill_set` = skills whose name is in `domain_tags`. If `|exposed_skill_set| < 2` after classification, fall back to the full installed set (flat exposure) and record `exposure_fallback: true` in the task board.545. Create session memory and task board with `exposed_skill_set` and `exposure_fallback` per task.5556### Scenes571. **PREPARE**: Plan, setup session ID, and initialize memory files.582. **ACT**: Spawn agents by priority tier within parallelism limits.593. **VERIFY**: Run self-check, `oma verify`, and QA cross-review loop.604. **RECOVER**: Retry failed agents with review history when limits allow.615. **FINALIZE**: Collect result files, compile summary, and clean progress files.6263### Transitions64- If native dispatch is available for current runtime/vendor, use it.65- If vendors differ or native path is unavailable, use fallback spawn.66- If verify or QA fails, feed feedback back to the implementation agent.67- If review loop limits are exceeded, report review history and quality warning.68- If a task's `exposed_skill_set` excludes a skill that a recovered failure indicates was needed, re-classify the task and re-dispatch with the expanded set rather than retrying against the original narrow set.6970### Failure and recovery71- Retry failed agents up to configured limits.72- Re-spawn with review history when review loop is exhausted.73- Pause or request re-specification when clarification debt thresholds are exceeded.7475### Exit76- Success: all tasks complete, verify/review pass, and results are summarized.77- Partial success: failed agents, exhausted review loops, or clarification debt are explicit.7879## Logical Operations8081### Actions82| Action | SSL primitive | Evidence |83|--------|---------------|----------|84| Read config and task context | `READ` | oma config, routing, request |85| Classify task into domain tags | `INFER` | task text vs each skill's `Intent signature` |86| Compute exposed skill set | `SELECT` | intersection of domain tags and installed skills |87| Select dispatch path | `SELECT` | Native vs fallback |88| Write session state | `WRITE` | task board and memory files |89| Spawn agents | `CALL_TOOL` | native CLI or `oma agent:spawn` |90| Poll progress | `READ` | progress/result files |91| Run verification | `CALL_TOOL` | `oma verify`, tests, QA |92| Update retry state | `UPDATE_STATE` | loop counters and CD metrics |93| Report final result | `NOTIFY` | compiled summary |9495### Tools and instruments96- Native CLI subagent dispatch, fallback spawn scripts, memory tools, verify script, QA agent97- Session metrics, prompt templates, task templates9899### Canonical command path100```bash101oma agent:spawn <agent-type> "<task>" <session-id> -w <workspace>102oma verify <agent-type> --workspace <workspace> --json103```104105When native runtime dispatch is available, prefer the runtime-specific native path listed in this skill before falling back to `oma agent:spawn`.106107### Resource scope108| Scope | Resource target |109|-------|-----------------|110| `LOCAL_FS` | Session, task-board, progress, result, config files |111| `PROCESS` | Agent CLI processes and verify scripts |112| `MEMORY` | Session state and clarification debt |113| `CODEBASE` | Workspaces owned by spawned agents |114115### Preconditions116- Task is decomposable into specialist agent work.117- Runtime/vendor dispatch path or fallback exists.118119### Effects and side effects120- Spawns agents and writes session/progress/result artifacts.121- May cause code changes through specialist agents.122- May trigger iterative review and retries.123124### Guardrails1251. Orchestrate per-agent dispatch from the project configuration before spawning any agent.1262. If `target_vendor === current_runtime_vendor` and the runtime has a verified native path, use native dispatch.1273. Otherwise fall back to `oma agent:spawn`.1284. Never exceed the configured parallelism or retry limits.1295. Keep session state, task-board state, progress files, and result files aligned throughout the run.1306. Domain gating must be soft: prefer a narrower `exposed_skill_set`, but fall back to flat exposure when classification confidence is low rather than starving a task of a required specialist.131132Current native executor paths:133- Claude Code: Agent tool with `.claude/agents/{agent}.md` definitions (multiple Agent tool calls in one message run in parallel; results return synchronously — no polling)134- OpenCode: native `task` tool with `subagent_type: {agent-id}`; do not use `oma agent:spawn` for same-session OpenCode work because it will not appear as a native child task135- Codex CLI: `codex exec "@agent ..."` using `.codex/agents/*.toml`136- Gemini CLI: `gemini -p "@agent ..."` using `.gemini/agents/*.md`137138Vendor-specific execution protocols are injected automatically for fallback CLI runs.139140### Configuration141142| Setting | Default | Description |143|---------|---------|-------------|144| MAX_PARALLEL | 3 | Max concurrent subagents |145| MAX_RETRIES | 2 | Retry attempts per failed task |146| POLL_INTERVAL | 30s | Status check interval |147| MAX_TURNS (impl) | 20 | Turn limit for backend/frontend/mobile |148| MAX_TURNS (review) | 15 | Turn limit for qa/debug |149| MAX_TURNS (plan) | 10 | Turn limit for pm |150151These are skill-level defaults applied by the orchestrating agent; they are not read from `config/cli-config.yaml` (which carries only vendor CLI and execution settings such as `results_dir` and `timeout`).152153### Memory Configuration154155Memory provider and tool names are configurable via `.agents/mcp.json` (not the repo-root `.mcp.json`, which is the Claude Code MCP server config):156```json157{158 "memoryConfig": {159 "provider": "file",160 "basePath": ".agents/state/memories",161 "tools": {162 "read": "Read",163 "write": "Write",164 "edit": "Edit"165 }166 }167}168```169170### Workflow Phases171172**PHASE 1 - Plan**: Analyze request -> decompose tasks -> generate session ID173**PHASE 1.5 - Domain gate**: For each task, intersect `Intent signature` matches across installed skills to derive `exposed_skill_set`. Record `exposure_fallback: true` when the intersection is too small to be useful and the flat library is used instead.174**PHASE 2 - Setup**: Use memory write tool to create `orchestrator-session.md` + `task-board.md` (include `exposed_skill_set` per task)175**PHASE 3 - Execute**: Spawn agents by priority tier (never exceed MAX_PARALLEL); inject only `exposed_skill_set` into each subagent's available specialist list176**PHASE 4 - Monitor**: Poll every POLL_INTERVAL; handle completed/failed/crashed agents177**PHASE 4.5 - Verify**: Run mechanical checks for every completed agent; run `oma verify {agent-type}` only for `backend`, `frontend`, `mobile`, `qa`, `debug`, and `pm`; then run QA cross-review for every completed implementation178**PHASE 5 - Collect**: Read all `result-{agent}-{sessionId}.md`, compile summary, cleanup progress files179180See `resources/subagent-prompt-template.md` for prompt construction.181See `resources/memory-schema.md` for memory file formats.182183### Memory File Ownership184185| File | Owner | Others |186|------|-------|--------|187| `orchestrator-session.md` | orchestrator | read-only |188| `task-board.md` | orchestrator | read-only |189| `progress-{agent}[-{sessionId}].md` | that agent | orchestrator reads |190| `result-{agent}[-{sessionId}].md` | that agent | orchestrator reads |191192### Agent-to-Agent Review Loop (PHASE 4.5)193194After each agent completes, enter an iterative review loop, not a single-pass verification.195196### Loop Flow197198```199Agent completes work200 ↓201[1] Mechanical Self-Check: lint, type-check, tests, diff scope202 ↓203[2] Verify: For supported types, run `oma verify {agent-type} --workspace {workspace}`204 Unsupported (`db`, `refactor`, `architecture`, `tf-infra`, `docs`) → record SKIP and continue205 ↓ FAIL → Agent receives feedback, fixes, back to [1]206 ↓ PASS207[3] Cross-Review: QA agent reviews the changes208 ↓ FAIL → Agent receives review feedback, fixes, back to [1]209 ↓ PASS210Accept result211```212213### Step Details214215**[1] Mechanical Self-Check** (formerly "Self-Review"):216Before requesting external review, the implementation agent must:217- Run lint, type-check, and tests in the workspace218- Verify only planned files were modified (diff scope check)219- Fix any mechanical failures (compile errors, test failures)220221**Quality judgment is NOT performed in this step.**222Design quality, architecture alignment, and acceptance criteria satisfaction223are evaluated exclusively in [3] Cross-Review by the QA agent.224Reason: Self-evaluation bias causes agents to consistently overrate their own output225(ref: Anthropic harness design research).226227**[2] Automated Verify**:228```bash229oma verify {agent-type} --workspace {workspace} --json230```231- Run only for `backend`, `frontend`, `mobile`, `qa`, `debug`, and `pm`.232- For `db`, `refactor`, `architecture`, `tf-infra`, and `docs`, record that automated verify is unsupported and continue to QA cross-review after the mechanical checks.233- **PASS (exit 0)**: Proceed to cross-review234- **FAIL (exit 1)**: Feed verify output back to the agent as correction context235236**[3] Cross-Review**: Spawn QA agent to review the changes:237- QA agent reads the diff, runs checks, evaluates against acceptance criteria238<!-- oma-docs:ignore-start -->239- If `docs/CODE-REVIEW.md` exists, QA agent uses it as the review checklist240<!-- oma-docs:ignore-end -->241- QA agent outputs: PASS (with optional nits) or FAIL (with specific issues)242- On FAIL: issues are fed back to the implementation agent for fixing243244### Loop Limits245246| Counter | Max | On Exceeded |247|---------|-----|-------------|248| Self-check + fix cycles | 3 | Escalate to cross-review regardless |249| Cross-review rejections | 2 | Report to user with review history |250| Total loop iterations | 5 | Force-complete with quality warning |251252### Review Feedback Format253254When feeding review results back to the implementation agent:255```256## Review Feedback (iteration {n}/{max})257**Reviewer**: {self / verify / qa-agent}258**Verdict**: FAIL259**Issues**:2601. {specific issue with file and line reference}2612. {specific issue}262**Fix instruction**: {what to change}263```264265This replaces single-pass verification. Most "nitpicking" should happen agent-to-agent.266Human review is reserved for final approval, not catching lint errors.267268### Retry Logic (after review loop exhaustion)269270Before starting any retry, check the termination conditions (OR, whichever fires first wins):2711. **Retry cap**: retry count for this agent has reached MAX_RETRIES — do not start another cycle.2722. **Session cost cap**: if a quota cap is configured (`loadQuotaCap()` from `cli/io/session-cost.ts`; no cap → skip), call `checkCap(sessionId, cap)`. On `exceeded === true`, save the agent's partial results, report early termination due to quota, and do not spawn the next retry or any remaining agents in the tier.273274If neither condition fires:275- 1st retry: Re-spawn agent with full review history as context276- 2nd retry: Re-spawn with "Try a different approach" + review history277- After MAX_RETRIES exhausted (cost cap not exceeded): activate the **Exploration Loop** (see `orchestrate.md` Step 5): generate 2-3 alternative hypotheses, spawn the same agent type with different hypothesis prompts in parallel separate workspaces, score with Quality Score when available, keep the highest-scoring approach, and record all experiments in the Experiment Ledger.278- Final failure: Report to user with complete review trail, ask whether to continue or abort279280### Clarification Debt (CD) Monitoring281282Track user corrections during session execution. See `../_shared/core/session-metrics.md` for full protocol.283284### Event Classification285When user sends feedback during session:286- **clarify** (+10): User answering agent's question287- **correct** (+25): User correcting agent's misunderstanding288- **redo** (+40): User rejecting work, requesting restart289290### Threshold Actions291| CD Score | Action |292|----------|--------|293| CD >= 50 | **RCA Required**: QA agent must add entry to `lessons-learned.md` |294| CD >= 80 | **Session Pause**: Request user to re-specify requirements |295| `redo` >= 2 | **Scope Lock**: Request explicit allowlist confirmation before continuing |296297### Recording298After each user correction event:299```300[EDIT]("session-metrics.md", append event to Events table)301```302303At session end, if CD >= 50:3041. Include CD summary in final report3052. Trigger QA agent RCA generation3063. Update `lessons-learned.md` with prevention measures307308309310## References311- Prompt template: `resources/subagent-prompt-template.md`312- Memory schema: `resources/memory-schema.md`313- Config: `config/cli-config.yaml`314- Scripts: `scripts/spawn-agent.sh`, `scripts/parallel-run.sh`, `scripts/verify.sh`315- Task templates: `templates/`316- Skill-to-agent mapping: `../_shared/core/skill-routing.md`317- Verification: `scripts/verify.sh <agent-type>`318- Session metrics: `../_shared/core/session-metrics.md`319- API contract template (SSOT): `../_shared/core/api-contracts/template.md`; read generated contracts from `.agents/results/api-contracts/` (run artifact) or `docs/plans/contracts/` (durable spec)320- Context loading: `../_shared/core/context-loading.md`321- Difficulty guide: `../_shared/core/difficulty-guide.md`322- Clarification protocol: `../_shared/core/clarification-protocol.md`323- Context budget: `../_shared/core/context-budget.md`324- Lessons learned: `../_shared/core/lessons-learned.md`