Multi-CLI Orchestrator Proposal (Claude/Codex/Gemini/OpenCode/Qwen)
1) Product Positioning
Build a neutral orchestration layer for coding agents.
- Any provider can be the entrypoint.
- Any provider can be a downstream reviewer.
- Provider participation is dynamic based on user environment and policy.
This avoids single-vendor lock-in and supports mixed-tool teams.
2) Provider Adapter Contract
Treat each CLI as a plugin adapter:
claude
codex
gemini
opencode
qwen
type ProviderId = "claude" | "codex" | "gemini" | "opencode" | "qwen";
interface CapabilitySet {
tiers: Array<"C0" | "C1" | "C2" | "C3" | "C4" | "C5" | "C6">;
supports_native_async: boolean;
supports_poll_endpoint: boolean;
supports_resume_after_restart: boolean;
supports_schema_enforcement: boolean;
min_supported_version: string;
tested_os: Array<"macos" | "linux" | "windows">;
}
interface TaskRunRef {
task_id: string;
provider: ProviderId;
run_id: string;
pid?: number; // shim async mode
session_id?: string; // native session mode
artifact_path: string;
started_at: string;
}
interface ProviderAdapter {
id: ProviderId;
detect(): Promise<ProviderPresence>;
capabilities(): Promise<CapabilitySet>;
run(input: TaskInput): Promise<TaskRunRef>;
poll(ref: TaskRunRef): Promise<TaskStatus>;
cancel(ref: TaskRunRef): Promise<void>;
normalize(raw: unknown, ctx: NormalizeContext): NormalizedFinding[];
}
3) Capability Tiering Standard
Use capability tiers instead of binary yes/no support.
C0: Binary + auth detection
C1: Non-interactive execution
C2: Structured JSON output (final result)
C3: Streaming structured events
C4: Session resume/continue
C5: Schema-constrained output
C6: Policy hooks/permissions/API control plane
Routing and SLOs should be based on these tiers, not provider names.
4) Initial Capability Matrix (Evidence-Based)
Y = supported in docs, P = partial/conditional (must pass probe before enabling), N = not documented.
| Provider |
C1 |
C2 |
C3 |
C4 |
C5 |
C6 |
Native async |
min_version (lock) |
tested_os |
Notes |
| Claude Code |
Y |
Y |
Y |
Y |
Y |
Y |
P |
2.1.59 (macos) |
macos |
Strong CLI/hook support; native job endpoint not primary mode |
| Codex CLI |
Y |
Y |
Y |
Y |
Y |
P |
P |
0.46.0 (macos) |
macos |
exec --json, resume; async mostly process/session based |
| Gemini CLI |
Y |
Y |
Y |
P |
N |
Y |
P |
0.1.7 (macos) |
macos |
Headless formats solid; resume semantics require probe |
| OpenCode |
Y |
Y |
Y |
Y |
N |
Y |
Y |
1.2.11 (macos) |
macos |
serve provides native pollable interface |
| Qwen Code |
Y |
Y |
P |
Y |
N |
P |
P |
0.10.6 (macos) |
macos |
stream-json output/input behavior varies by version; C3 passed with qwen-oauth auth mode |
Implementation rule:
- If task needs only
C1-C2, all qualifying providers can run.
- If task requires native async polling, prefer providers with
supports_native_async=true, else use shim polling.
Probe pass criteria (frozen before implementation):
C0: binary found, version readable, auth check successful.
C1: non-interactive fixture run yields expected completion token/output.
C2: final output includes machine-parseable JSON content that passes shape validation.
C3: stream mode emits >= 2 valid structured events on fixture task.
C4: resume across fresh process works and context continuity is verified.
C5: schema-constrained run rejects invalid schema and respects valid schema.
C6: permission/policy constraints enforced in negative tests.
5) Runtime Discovery and User Config
Discovery order
- Load user config (
enabled, roles, allowlist).
- Probe binaries (
which claude, which codex, etc.).
- Run auth checks and capability probes.
- Register healthy adapters with measured capabilities.
Config example
providers:
claude:
enabled: true
role: [entrypoint, reviewer]
weight: 1.0
max_cost_usd: 0.8
min_version: "2.1.59"
tested_os: [macos]
codex:
enabled: true
role: [entrypoint, reviewer]
weight: 1.0
max_cost_usd: 0.8
min_version: "0.46.0"
tested_os: [macos]
gemini:
enabled: true
role: [reviewer]
weight: 0.8
max_cost_usd: 0.6
min_version: "0.1.7"
tested_os: [macos]
opencode:
enabled: true
role: [reviewer]
weight: 0.7
max_cost_usd: 0.5
min_version: "1.2.11"
tested_os: [macos]
qwen:
enabled: true
role: [reviewer]
weight: 0.6
max_cost_usd: 0.5
min_version: "0.10.6"
tested_os: [macos]
policy:
required_capabilities: [C1, C2]
optional_capabilities: [C3, C4, C5, C6]
max_parallel_reviewers: 2
budget_usd_per_task: 1.5
timeout_seconds: 600
provider_allowlist: [claude, codex, gemini, opencode, qwen]
6) Task State Machine and Idempotency
Global task state
DRAFT -> QUEUED -> DISPATCHED -> RUNNING -> AGGREGATING -> COMPLETED
- Retry branch:
RUNNING -> RETRYING -> RUNNING
- Terminal states:
COMPLETED, PARTIAL_SUCCESS, FAILED, CANCELLED, EXPIRED
Provider attempt state
PENDING -> STARTED -> SUCCEEDED
PENDING -> STARTED -> RETRYABLE_FAILED -> RETRYING -> STARTED
PENDING -> STARTED -> NON_RETRYABLE_FAILED
PENDING -> STARTED -> CANCELLED
Idempotency model
task_idempotency_key = hash(repo, revision, prompt, scope, policy_version)
- Same key returns existing active task instead of creating a duplicate.
dispatch_key = hash(task_id, provider, attempt_no) prevents duplicate fan-out.
- Artifact writes are atomic (
tmp then rename).
- Notifications are deduped by
(task_id, terminal_state, channel).
Partial success policy
- If at least one required provider succeeds and policy gates pass:
PARTIAL_SUCCESS.
- If all required providers fail:
FAILED.
EXPIRED trigger and reaper
- Expire condition A: run wall-clock exceeds
timeout_seconds + grace_seconds.
- Expire condition B: no heartbeat update for
heartbeat_ttl_seconds.
- Reaper scan interval: every 60 seconds (configurable).
- Reaper compensation flow:
- Try provider-native cancel if available.
- If shim mode, send
SIGTERM, wait 10 seconds, then SIGKILL.
- Mark attempt/task
EXPIRED, persist partial artifacts, emit terminal notification.
- Apply retry policy only if attempt is marked retryable by adapter error classifier.
7) Async Execution Model and Poll Semantics
poll(ref) must work in both modes:
- Native async mode (
supports_native_async=true)
- Adapter submits run and gets provider-side
run_id/session_id.
poll() queries provider endpoint/session status.
- Shim async mode (
supports_native_async=false)
- Orchestrator starts CLI as OS process and tracks
pid.
- Stdout/stderr are streamed to artifact files.
poll() checks process state + heartbeat file + output completion marker.
- On process exit, adapter parses final artifact and maps terminal state.
Shim guarantees:
- Works for CLIs that are synchronous by design.
- Supports timeout, cancellation, and crash recovery via persisted run refs.
8) Capability-Gated Routing
Hard filters:
- Provider is enabled and allowlisted.
- Provider meets
required_capabilities.
- Provider has not exceeded per-provider cost cap.
- Estimated run cost <= remaining task budget.
Scoring formula:
score = capability_match * reliability_weight * user_weight / expected_cost
Selection defaults:
- Low risk diff: top 1 reviewer
- Medium risk diff: top 2 reviewers
- High risk/security-tagged diff: top 3 reviewers (budget permitting)
9) Normalize Strategy (Critical Path)
normalize() is implementation-critical and provider-specific.
| Provider type |
Strategy |
Failure fallback |
| Structured-output first (Claude/Codex when schema enabled) |
Use JSON/schema mode, direct field mapping to canonical schema |
If schema parse fails, rerun with strict JSON-only prompt once |
| JSON-capable but schema-weak (Gemini/OpenCode/Qwen in some modes) |
Use structured prompt contract + JSON parser + repair pass |
If repair fails, mark run as parse-failed and trigger text-parser fallback |
| Text-first outputs |
Two-step extraction: text -> LLM extractor constrained to canonical schema |
If extractor fails twice, keep raw artifact and emit normalization_error |
Prompt contract for extraction mode:
- Require array of findings with fixed keys.
- Require file path and line when available (
line can be null).
- For unknown fields, return
null instead of inventing values.
Normalization quality checks:
- Drop findings without category/severity/title.
- Clamp confidence to
[0,1].
- Validate file paths are inside repo root.
10) Canonical Finding Schema
{
"task_id": "string",
"provider": "claude|codex|gemini|opencode|qwen",
"finding_id": "string",
"severity": "critical|high|medium|low",
"category": "bug|security|performance|maintainability|test-gap",
"title": "string",
"evidence": {
"file": "path",
"line": null,
"symbol": "optional symbol",
"snippet": "string"
},
"recommendation": "string",
"confidence": 0.0,
"fingerprint": "stable hash",
"raw_ref": "artifact pointer"
}
Schema note:
evidence.line is nullable for providers that only return symbol/snippet-level evidence.
11) Dedupe and Conflict Resolution
file+line+category is stage-0 only. Use three-stage merge:
- Exact fingerprint merge
- Normalized path + symbol + category + canonicalized title hash.
- Proximity merge
- Same file + same category + line delta <= 5 + token similarity above threshold.
- Semantic merge
- Embedding similarity above threshold and overlapping evidence snippets.
Conflict rules:
- Severity = max severity in cluster.
- Confidence = weighted average by provider reliability.
- Preserve per-provider evidence links for audit.
12) Security and Compliance Baseline
Data egress controls
- Provider allowlist and per-project egress policy.
- Pre-send redaction for secrets/tokens/credentials/PII.
- Path filters (
include_paths/exclude_paths) to minimize code exposure.
Access controls
- Default read-only tool profile for review tasks.
- Provider-specific permission profiles; no implicit shell escalation.
- Global kill switch per provider.
Audit and retention
- Immutable attempt log: actor, timestamps, provider, prompt hash, artifact hashes.
- Retention by class (
raw_artifacts, summaries, decision_log).
- Traceability from merged findings to raw provider outputs.
Compliance hooks
- Optional policy checks before dispatch (license/classification/region rules).
- Block dispatch if project policy requires local-only and provider is remote-only.
13) Artifact Generation Logic (summary.md / decision.md)
Frozen stage-A contract implementation:
- Adapter interface:
runtime/contracts.py
- Artifact layout:
runtime/artifacts.py
summary.md generation:
- Rule-based renderer from canonical findings.
- Includes counts by severity/category, provider coverage, and unresolved parse errors.
decision.md generation:
- Phase 1 (MVP): rule-based decision engine only.
- Decision rules: fail gate on
critical, escalate on high >= threshold, else pass with follow-ups.
- Phase 2: optional AI-assisted synthesis pass (provider configurable), but final decision must include rule trace.
This keeps decision.md deterministic in early stages and auditable later.
Stage-A required per-task artifacts:
summary.md
decision.md
findings.json
run.json
providers/<provider>.json
raw/<provider>.stdout.log
raw/<provider>.stderr.log
14) Risk Register and Mitigations
| Risk |
Impact |
Mitigation |
| Output format divergence |
Normalization failures |
Adapter-specific normalize strategy + parse fallback + error taxonomy |
| Auth/config variation |
Provider unavailable at runtime |
doctor checks + setup guide + preflight validation |
| Cost overrun |
Budget breach |
Global budget + per-provider cap + auto fan-out downgrade |
| False dedupe or missed merge |
Quality drop |
3-stage dedupe + sampled human audit + threshold tuning |
| Async instability |
Orphaned runs |
Native/shim poll model + heartbeat + timeout + restart recovery |
15) Realistic Delivery Plan
Stage A (2 weeks): gate + core
- State machine, idempotency, artifact writer, notifier
- Adapter framework + capability probes
- Two providers only:
claude, codex
Stage B (2-3 weeks): expansion + security
- Add
gemini, opencode
- Redaction, allowlist, audit logs
- Dedupe v2 (exact + proximity)
Stage C (2 weeks): fifth provider + optimization
- Add
qwen
- Semantic dedupe and conflict scoring
- Metrics dashboard + routing calibration
16) Implementation Gate Checklist
Do not start broad rollout until all pass, in this order:
- Freeze capability probe pass/fail criteria for each
C level and lock them in test fixtures.
- Run adapter contract tests: input/output mapping, error taxonomy, retry semantics.
- Run 2-provider dry run (
claude + codex) to validate idempotency, state transitions, and artifact persistence chain.
- Run state-machine suite (retries, duplicate submit, partial success, expired recovery).
- Verify security baseline (redaction, allowlist, audit log).
- Verify budget/latency policies in target environment.
17) Success Criteria by Stage
Stage A (leading indicators):
= 90% successful normalization on sampled runs
- zero duplicate notifications for duplicate submissions
- median latency for 1-provider review within target SLO
Stage B (operational quality):
- parse-failure rate < 5%
- budget overrun rate < 3%
- manual reviewer "useful" rating >= 4/5 on pilot tasks
Stage C (outcome metrics, 4-8 week pilot):
= 20% increase in accepted high-severity findings
- <= 15% increase in median feedback time
- measurable false-positive reduction after dedupe tuning
- < 3% task failure from adapter/runtime issues
If metrics miss targets, reduce fan-out depth and keep best-performing providers.
18) Sources and Known Uncertainties
Primary sources:
- Claude Code CLI and hooks docs
- OpenAI Codex CLI
exec/non-interactive docs
- Gemini CLI headless and command argument docs
- OpenCode CLI and agents docs
- Qwen Code non-interactive/resume/output docs
Known uncertainty:
- Gemini headless resume behavior across process restarts is treated as partial until adapter probe confirms.
1---2name: multi-cli-orchestrator-proposal-claude-codex-gemini3description: Routing and SLOs should be based on these tiers, not provider names.4---5# Multi-CLI Orchestrator Proposal (Claude/Codex/Gemini/OpenCode/Qwen)67## 1) Product Positioning8Build a neutral orchestration layer for coding agents.910- Any provider can be the entrypoint.11- Any provider can be a downstream reviewer.12- Provider participation is dynamic based on user environment and policy.1314This avoids single-vendor lock-in and supports mixed-tool teams.1516## 2) Provider Adapter Contract17Treat each CLI as a plugin adapter:1819- `claude`20- `codex`21- `gemini`22- `opencode`23- `qwen`2425```ts26type ProviderId = "claude" | "codex" | "gemini" | "opencode" | "qwen";2728interface CapabilitySet {29 tiers: Array<"C0" | "C1" | "C2" | "C3" | "C4" | "C5" | "C6">;30 supports_native_async: boolean;31 supports_poll_endpoint: boolean;32 supports_resume_after_restart: boolean;33 supports_schema_enforcement: boolean;34 min_supported_version: string;35 tested_os: Array<"macos" | "linux" | "windows">;36}3738interface TaskRunRef {39 task_id: string;40 provider: ProviderId;41 run_id: string;42 pid?: number; // shim async mode43 session_id?: string; // native session mode44 artifact_path: string;45 started_at: string;46}4748interface ProviderAdapter {49 id: ProviderId;50 detect(): Promise<ProviderPresence>;51 capabilities(): Promise<CapabilitySet>;52 run(input: TaskInput): Promise<TaskRunRef>;53 poll(ref: TaskRunRef): Promise<TaskStatus>;54 cancel(ref: TaskRunRef): Promise<void>;55 normalize(raw: unknown, ctx: NormalizeContext): NormalizedFinding[];56}57```5859## 3) Capability Tiering Standard60Use capability tiers instead of binary yes/no support.6162- `C0`: Binary + auth detection63- `C1`: Non-interactive execution64- `C2`: Structured JSON output (final result)65- `C3`: Streaming structured events66- `C4`: Session resume/continue67- `C5`: Schema-constrained output68- `C6`: Policy hooks/permissions/API control plane6970Routing and SLOs should be based on these tiers, not provider names.7172## 4) Initial Capability Matrix (Evidence-Based)73`Y` = supported in docs, `P` = partial/conditional (must pass probe before enabling), `N` = not documented.7475| Provider | C1 | C2 | C3 | C4 | C5 | C6 | Native async | min_version (lock) | tested_os | Notes |76| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |77| Claude Code | Y | Y | Y | Y | Y | Y | P | 2.1.59 (macos) | macos | Strong CLI/hook support; native job endpoint not primary mode |78| Codex CLI | Y | Y | Y | Y | Y | P | P | 0.46.0 (macos) | macos | `exec --json`, `resume`; async mostly process/session based |79| Gemini CLI | Y | Y | Y | P | N | Y | P | 0.1.7 (macos) | macos | Headless formats solid; resume semantics require probe |80| OpenCode | Y | Y | Y | Y | N | Y | Y | 1.2.11 (macos) | macos | `serve` provides native pollable interface |81| Qwen Code | Y | Y | P | Y | N | P | P | 0.10.6 (macos) | macos | `stream-json` output/input behavior varies by version; C3 passed with `qwen-oauth` auth mode |8283Implementation rule:84- If task needs only `C1-C2`, all qualifying providers can run.85- If task requires native async polling, prefer providers with `supports_native_async=true`, else use shim polling.8687Probe pass criteria (frozen before implementation):88- `C0`: binary found, version readable, auth check successful.89- `C1`: non-interactive fixture run yields expected completion token/output.90- `C2`: final output includes machine-parseable JSON content that passes shape validation.91- `C3`: stream mode emits >= 2 valid structured events on fixture task.92- `C4`: resume across fresh process works and context continuity is verified.93- `C5`: schema-constrained run rejects invalid schema and respects valid schema.94- `C6`: permission/policy constraints enforced in negative tests.9596## 5) Runtime Discovery and User Config97### Discovery order981. Load user config (`enabled`, roles, allowlist).992. Probe binaries (`which claude`, `which codex`, etc.).1003. Run auth checks and capability probes.1014. Register healthy adapters with measured capabilities.102103### Config example104```yaml105providers:106 claude:107 enabled: true108 role: [entrypoint, reviewer]109 weight: 1.0110 max_cost_usd: 0.8111 min_version: "2.1.59"112 tested_os: [macos]113 codex:114 enabled: true115 role: [entrypoint, reviewer]116 weight: 1.0117 max_cost_usd: 0.8118 min_version: "0.46.0"119 tested_os: [macos]120 gemini:121 enabled: true122 role: [reviewer]123 weight: 0.8124 max_cost_usd: 0.6125 min_version: "0.1.7"126 tested_os: [macos]127 opencode:128 enabled: true129 role: [reviewer]130 weight: 0.7131 max_cost_usd: 0.5132 min_version: "1.2.11"133 tested_os: [macos]134 qwen:135 enabled: true136 role: [reviewer]137 weight: 0.6138 max_cost_usd: 0.5139 min_version: "0.10.6"140 tested_os: [macos]141142policy:143 required_capabilities: [C1, C2]144 optional_capabilities: [C3, C4, C5, C6]145 max_parallel_reviewers: 2146 budget_usd_per_task: 1.5147 timeout_seconds: 600148 provider_allowlist: [claude, codex, gemini, opencode, qwen]149```150151## 6) Task State Machine and Idempotency152### Global task state153- `DRAFT` -> `QUEUED` -> `DISPATCHED` -> `RUNNING` -> `AGGREGATING` -> `COMPLETED`154- Retry branch: `RUNNING` -> `RETRYING` -> `RUNNING`155- Terminal states: `COMPLETED`, `PARTIAL_SUCCESS`, `FAILED`, `CANCELLED`, `EXPIRED`156157### Provider attempt state158- `PENDING` -> `STARTED` -> `SUCCEEDED`159- `PENDING` -> `STARTED` -> `RETRYABLE_FAILED` -> `RETRYING` -> `STARTED`160- `PENDING` -> `STARTED` -> `NON_RETRYABLE_FAILED`161- `PENDING` -> `STARTED` -> `CANCELLED`162163### Idempotency model164- `task_idempotency_key = hash(repo, revision, prompt, scope, policy_version)`165- Same key returns existing active task instead of creating a duplicate.166- `dispatch_key = hash(task_id, provider, attempt_no)` prevents duplicate fan-out.167- Artifact writes are atomic (`tmp` then rename).168- Notifications are deduped by `(task_id, terminal_state, channel)`.169170### Partial success policy171- If at least one required provider succeeds and policy gates pass: `PARTIAL_SUCCESS`.172- If all required providers fail: `FAILED`.173174### `EXPIRED` trigger and reaper175- Expire condition A: run wall-clock exceeds `timeout_seconds + grace_seconds`.176- Expire condition B: no heartbeat update for `heartbeat_ttl_seconds`.177- Reaper scan interval: every 60 seconds (configurable).178- Reaper compensation flow:1791. Try provider-native cancel if available.1802. If shim mode, send `SIGTERM`, wait 10 seconds, then `SIGKILL`.1813. Mark attempt/task `EXPIRED`, persist partial artifacts, emit terminal notification.1824. Apply retry policy only if attempt is marked retryable by adapter error classifier.183184## 7) Async Execution Model and Poll Semantics185`poll(ref)` must work in both modes:1861871. Native async mode (`supports_native_async=true`)188- Adapter submits run and gets provider-side `run_id/session_id`.189- `poll()` queries provider endpoint/session status.1901912. Shim async mode (`supports_native_async=false`)192- Orchestrator starts CLI as OS process and tracks `pid`.193- Stdout/stderr are streamed to artifact files.194- `poll()` checks process state + heartbeat file + output completion marker.195- On process exit, adapter parses final artifact and maps terminal state.196197Shim guarantees:198- Works for CLIs that are synchronous by design.199- Supports timeout, cancellation, and crash recovery via persisted run refs.200201## 8) Capability-Gated Routing202Hard filters:203- Provider is enabled and allowlisted.204- Provider meets `required_capabilities`.205- Provider has not exceeded per-provider cost cap.206- Estimated run cost <= remaining task budget.207208Scoring formula:209`score = capability_match * reliability_weight * user_weight / expected_cost`210211Selection defaults:212- Low risk diff: top 1 reviewer213- Medium risk diff: top 2 reviewers214- High risk/security-tagged diff: top 3 reviewers (budget permitting)215216## 9) Normalize Strategy (Critical Path)217`normalize()` is implementation-critical and provider-specific.218219| Provider type | Strategy | Failure fallback |220| --- | --- | --- |221| Structured-output first (Claude/Codex when schema enabled) | Use JSON/schema mode, direct field mapping to canonical schema | If schema parse fails, rerun with strict JSON-only prompt once |222| JSON-capable but schema-weak (Gemini/OpenCode/Qwen in some modes) | Use structured prompt contract + JSON parser + repair pass | If repair fails, mark run as parse-failed and trigger text-parser fallback |223| Text-first outputs | Two-step extraction: text -> LLM extractor constrained to canonical schema | If extractor fails twice, keep raw artifact and emit `normalization_error` |224225Prompt contract for extraction mode:226- Require array of findings with fixed keys.227- Require file path and line when available (`line` can be `null`).228- For unknown fields, return `null` instead of inventing values.229230Normalization quality checks:231- Drop findings without category/severity/title.232- Clamp confidence to `[0,1]`.233- Validate file paths are inside repo root.234235## 10) Canonical Finding Schema236```json237{238 "task_id": "string",239 "provider": "claude|codex|gemini|opencode|qwen",240 "finding_id": "string",241 "severity": "critical|high|medium|low",242 "category": "bug|security|performance|maintainability|test-gap",243 "title": "string",244 "evidence": {245 "file": "path",246 "line": null,247 "symbol": "optional symbol",248 "snippet": "string"249 },250 "recommendation": "string",251 "confidence": 0.0,252 "fingerprint": "stable hash",253 "raw_ref": "artifact pointer"254}255```256257Schema note:258- `evidence.line` is nullable for providers that only return symbol/snippet-level evidence.259260## 11) Dedupe and Conflict Resolution261`file+line+category` is stage-0 only. Use three-stage merge:2622631. Exact fingerprint merge264- Normalized path + symbol + category + canonicalized title hash.2652662. Proximity merge267- Same file + same category + line delta <= 5 + token similarity above threshold.2682693. Semantic merge270- Embedding similarity above threshold and overlapping evidence snippets.271272Conflict rules:273- Severity = max severity in cluster.274- Confidence = weighted average by provider reliability.275- Preserve per-provider evidence links for audit.276277## 12) Security and Compliance Baseline278### Data egress controls279- Provider allowlist and per-project egress policy.280- Pre-send redaction for secrets/tokens/credentials/PII.281- Path filters (`include_paths`/`exclude_paths`) to minimize code exposure.282283### Access controls284- Default read-only tool profile for review tasks.285- Provider-specific permission profiles; no implicit shell escalation.286- Global kill switch per provider.287288### Audit and retention289- Immutable attempt log: actor, timestamps, provider, prompt hash, artifact hashes.290- Retention by class (`raw_artifacts`, `summaries`, `decision_log`).291- Traceability from merged findings to raw provider outputs.292293### Compliance hooks294- Optional policy checks before dispatch (license/classification/region rules).295- Block dispatch if project policy requires local-only and provider is remote-only.296297## 13) Artifact Generation Logic (summary.md / decision.md)298Frozen stage-A contract implementation:299- Adapter interface: `runtime/contracts.py`300- Artifact layout: `runtime/artifacts.py`301302`summary.md` generation:303- Rule-based renderer from canonical findings.304- Includes counts by severity/category, provider coverage, and unresolved parse errors.305306`decision.md` generation:307- Phase 1 (MVP): rule-based decision engine only.308- Decision rules: fail gate on `critical`, escalate on `high >= threshold`, else pass with follow-ups.309- Phase 2: optional AI-assisted synthesis pass (provider configurable), but final decision must include rule trace.310311This keeps `decision.md` deterministic in early stages and auditable later.312313Stage-A required per-task artifacts:314- `summary.md`315- `decision.md`316- `findings.json`317- `run.json`318- `providers/<provider>.json`319- `raw/<provider>.stdout.log`320- `raw/<provider>.stderr.log`321322## 14) Risk Register and Mitigations323| Risk | Impact | Mitigation |324| --- | --- | --- |325| Output format divergence | Normalization failures | Adapter-specific normalize strategy + parse fallback + error taxonomy |326| Auth/config variation | Provider unavailable at runtime | `doctor` checks + setup guide + preflight validation |327| Cost overrun | Budget breach | Global budget + per-provider cap + auto fan-out downgrade |328| False dedupe or missed merge | Quality drop | 3-stage dedupe + sampled human audit + threshold tuning |329| Async instability | Orphaned runs | Native/shim poll model + heartbeat + timeout + restart recovery |330331## 15) Realistic Delivery Plan332### Stage A (2 weeks): gate + core333- State machine, idempotency, artifact writer, notifier334- Adapter framework + capability probes335- Two providers only: `claude`, `codex`336337### Stage B (2-3 weeks): expansion + security338- Add `gemini`, `opencode`339- Redaction, allowlist, audit logs340- Dedupe v2 (exact + proximity)341342### Stage C (2 weeks): fifth provider + optimization343- Add `qwen`344- Semantic dedupe and conflict scoring345- Metrics dashboard + routing calibration346347## 16) Implementation Gate Checklist348Do not start broad rollout until all pass, in this order:3493501. Freeze capability probe pass/fail criteria for each `C` level and lock them in test fixtures.3512. Run adapter contract tests: input/output mapping, error taxonomy, retry semantics.3523. Run 2-provider dry run (`claude` + `codex`) to validate idempotency, state transitions, and artifact persistence chain.3534. Run state-machine suite (retries, duplicate submit, partial success, expired recovery).3545. Verify security baseline (redaction, allowlist, audit log).3556. Verify budget/latency policies in target environment.356357## 17) Success Criteria by Stage358Stage A (leading indicators):359- >= 90% successful normalization on sampled runs360- zero duplicate notifications for duplicate submissions361- median latency for 1-provider review within target SLO362363Stage B (operational quality):364- parse-failure rate < 5%365- budget overrun rate < 3%366- manual reviewer "useful" rating >= 4/5 on pilot tasks367368Stage C (outcome metrics, 4-8 week pilot):369- >= 20% increase in accepted high-severity findings370- <= 15% increase in median feedback time371- measurable false-positive reduction after dedupe tuning372- < 3% task failure from adapter/runtime issues373374If metrics miss targets, reduce fan-out depth and keep best-performing providers.375376## 18) Sources and Known Uncertainties377Primary sources:378- Claude Code CLI and hooks docs379- OpenAI Codex CLI `exec`/non-interactive docs380- Gemini CLI headless and command argument docs381- OpenCode CLI and agents docs382- Qwen Code non-interactive/resume/output docs383384Known uncertainty:385- Gemini headless resume behavior across process restarts is treated as partial until adapter probe confirms.