zuvo:benchmark — Multi-Provider Coding Benchmark
Measures how well different AI coding agents handle a task. Dispatches to all available providers in parallel, collects responses, and uses an opposite-model Claude meta-judge to score on four dimensions (completeness, accuracy, actionability, no_hallucinations). Outputs a ranked leaderboard with cost and time breakdown.
Use corpus mode (--mode corpus) to compare runs across time. The corpus uses fixed OrderService + useSearchProducts tasks — the same prompt every time — so runs from different days are directly comparable.
Argument Parsing
Parse $ARGUMENTS for these flags:
| Flag | Effect |
|---|---|
--diff [ref] |
Benchmark against a git diff (default: HEAD~1) |
--files <path> |
Benchmark on a file or newline-separated list of files as task input |
--prompt <text> |
Use a literal text prompt as the task (alias: --task) |
--mode corpus |
Use fixed corpus tasks (OrderService + useSearchProducts) |
--mode default |
Use user-provided task (default) |
--with-tests |
Run Round 2: have providers write tests for their own Round 1 code |
--with-adversarial |
Run adversarial cross-review on Round 1 code |
--with-test-adversarial |
Run adversarial cross-review on Round 3 tests (requires --with-tests) |
--with-static-checks |
Run tsc + jest on generated code (best-effort, null if tools missing) |
--provider <name> |
Restrict to one or more providers, comma-separated (alias: --providers) |
--show-costs |
Print provider cost table ($/M tokens) and exit |
--compare [id1] [id2] |
Compare two prior runs from zuvo/reports/; default: last two (skill-only, not passed to runner) |
--replay-last |
Re-run benchmark with the same task as the most recent run (skill-only, not passed to runner) |
--json |
Output raw JSON to stdout instead of formatted leaderboard (handled by both runner and skill) |
--no-snapshot |
Suppress task_snapshot storage (use if task contains secrets or PII) |
--dry-run |
Print prompt + providers without dispatching |
| (remaining text) | Treated as --prompt <text> if no other input flag given |
Mandatory File Loading
Before starting work, read each file below. Print the checklist with status.
CORE FILES LOADED:
1. ../../shared/includes/benchmark-output-schema.md -- READ/MISSING
2. ../../shared/includes/run-logger.md -- READ/MISSING
3. ../../shared/includes/retrospective.md -- READ/MISSING
4. ../../shared/includes/env-compat.md -- READ/MISSING
If any core file is missing, proceed in degraded mode and note it in the BENCHMARK COMPLETE block.
Phase 0: Parse and Validate Arguments
Parse all flags from
$ARGUMENTSper the table above.--show-costs: runscripts/benchmark.sh --show-costs, print the table, stop. No benchmark run.--compare: load the two specified run JSONs fromzuvo/reports/(or the last two if no IDs given). Print a delta table: quality, time_s, cost_usd per provider. Warn if task_hash values differ (different tasks, comparison may be misleading). Stop — no benchmark run.--replay-last: find the most recent.jsonfile inzuvo/reports/. Readtask_snapshotas the task input. SetINPUT_MODE=prompt. Error if no history found. Error if--mode corpus(corpus uses fixed prompts, replay makes no sense).
Dispatch follows ../../shared/includes/execution-policy.md through env-compat. Reuse existing
authorization within that policy; session restrictions take precedence. Run each required gate
and report its actual independence or an unmet requirement.
Validate flag combinations:
--with-testsand--with-adversarialrequire--mode corpus(they assume a two-round structure). If set without--mode corpus, print a warning and auto-set--mode corpus.--providerwith a name not in the detected provider list → warn, but don't fail (provider may appear after dispatch attempt).
Determine task source:
--mode corpus→task_source: "corpus", load from../../shared/includes/benchmark-corpus/task-code.md--diff→task_source: "diff", input via git diff--files→task_source: "files", input via file concatenation--promptor remaining text →task_source: "user", input is the literal text- None of the above →
task_source: "diff", default toHEAD~1(matches runner behavior)
Print run header:
BENCHMARK RUN Mode: [default|corpus] | Task: [first 60 chars or "corpus"] | Providers: [list] Options: tests=[bool] adversarial=[bool] static=[bool]
Phase 1: Dispatch Providers
Find and run
benchmark.sh. Search in order:scripts/benchmark.sh(if in zuvo-plugin repo)~/.codex/scripts/benchmark.sh(Codex install)~/.cursor/scripts/benchmark.sh(Cursor install)~/.claude/plugins/cache/zuvo-marketplace/zuvo/*/scripts/benchmark.sh(Claude Code plugin cache)
Capture stdout to
$TMPDIR/benchmark-raw.json.scripts/benchmark.sh \ --mode <mode> \ [--prompt <text>|--files <path>|--diff <ref>] \ [--with-tests] [--with-adversarial] [--with-test-adversarial] [--with-static-checks] \ [--provider <list>] [--json] \ --run-id <run_id> \ --round-dir "$TMPDIR/rounds" \ --output "$TMPDIR/benchmark-raw.json"Runner-only flags:
--prompt,--files,--diff,--mode,--with-*,--provider,--output,--run-id,--round-dir,--json,--show-costs,--dry-run,--no-snapshot. Skill-only flags (NOT passed to runner):--compare,--replay-last. These are handled in Phase 0 before the runner is invoked.Check exit code:
0: success, proceed to Phase 21: bad arguments → show error and stop2: no providers available → show install instructions and stop3: all providers failed → show error and stop
Print progress as providers complete (read from stderr in real time or wait for script exit):
[1/4] claude — done (42s) [2/4] gemini — done (31s) [3/4] codex-fast — done (28s) [4/4] cursor-agent — timeout
Phase 2: Meta-Judge Scoring
Never skip this phase. Even if only one provider succeeded, still score it.
Model Selection (opposite-model rule)
The judge must be a different model than the one running this skill to reduce self-serving bias:
- If
$CLAUDE_MODELcontains "opus" → useclaude-sonnet-5 - Otherwise (sonnet, haiku, or unset) → use
claude-opus-5
This is the opposite model from the one executing this skill.
Input Preparation (CQ6 — bounded input)
For each provider response in providers_raw:
- Extract
response_excerptor the full response text - CQ6 truncation: limit each response to
floor(80000 / provider_count)characters before including in the judge prompt. Setjudge_input_truncated: truein output if any response was truncated.
Shuffle the presentation order randomly (store in judge_presentation_order) to reduce positional bias.
Judge Prompt
Send a single Claude call to the judge model with this prompt structure:
You are an objective code quality judge. Score each provider's response on four dimensions (0–5 each):
- completeness: All required methods, fields, and behaviors implemented
- accuracy: Logic is correct, state machines valid, edge cases handled
- actionability: Production-ready; no stubs, TODOs, or placeholder logic
- no_hallucinations: No invented APIs, non-existent methods, or fabricated behavior
TASK:
<task_snapshot or task prompt>
PROVIDER RESPONSES (in random order):
[provider ID redacted — labeled A, B, C, D]
<response A>
---
<response B>
...
Respond with JSON only, no prose:
{
"A": { "completeness": N, "accuracy": N, "actionability": N, "no_hallucinations": N },
"B": { ... },
...
}
Parse Judge Response
- Extract JSON from the response. If
jq .fails → mark runUNSCORED, skip to Phase 3 with null scores. Log the raw judge output for debugging. - Map labeled responses (A, B, C...) back to provider names using
judge_presentation_order. - Compute per-provider:
code_composite = completeness + accuracy + actionability + no_hallucinations(0–20)quality = round(code_composite * 5)(0–100) — code-only modeself_eval_bias = self_eval_raw - code_composite— null ifself_eval_rawis null
Phase 3: Assemble Leaderboard
Rank Providers
Sort by:
qualityDESC (higher is better)time_sASC (faster wins ties)cost_usdASC (cheaper wins cost ties)provideralphabetical (deterministic final tiebreaker)
Build Leaderboard Array
For each provider in ranked order:
{
"rank": 1,
"provider": "claude",
"quality": 87,
"code_score": 18,
"test_score": null,
"time_s": 42.1,
"cost_usd": 0.031,
"compile_ok": null,
"tests_pass": null,
"self_eval_bias": 1.2,
"adversarial_delta": null,
"status": "scored"
}
Build Scorecards Object
{
"claude": {
"code_completeness": 5,
"code_accuracy": 4,
"code_actionability": 5,
"code_no_hallucinations": 4,
"code_composite": 18,
"test_completeness": null,
"test_accuracy": null,
"test_actionability": null,
"test_no_hallucinations": null,
"test_composite": null,
"adversarial_delta": null,
"test_adversarial_delta": null,
"self_eval_bias": 1.2,
"response_excerpt": "<first 500 chars of response>"
}
}
Leaderboard Display
Print the leaderboard as a markdown table:
## Benchmark Results — [run_id]
Task: [first 60 chars] | Mode: [default|corpus]
| Rank | Provider | Quality | Code | Tests | Time | Cost | Tokens | Adv.Δ | Status |
|------|-------------|---------|------|-------|-------|---------|------------|-------|--------|
| 1 | codex-fast | 85 | 17 | 17 | 345s | $0.025 | ~12K/~8K | -2 | scored |
| 2 | gemini | 75 | 15 | 15 | 291s | $0.000 | ~11K/~9K | -4 | scored |
| 3 | claude | 70 | 14 | 14 | 300s | $0.062 | ~12K/~7K | -3 | scored |
| 4 | cursor-agent | 65 | 13 | 13 | 357s | — | ~11K/~6K | -3 | scored |
Tokens column: input/output estimates (~ = estimated via word count × 1.3).
Self-eval bias (positive = overconfident): codex-fast +1, claude +4, cursor-agent +5, gemini +5
Adversarial delta (negative = score dropped after adversarial critique): codex-fast -2, claude -3, cursor-agent -3, gemini -4
Phase 4: Persist and Log
Output Files
Auto-increment the run number (NNN) by reading existing files in zuvo/reports/:
zuvo/reports/benchmark-NNN-[run_id].md— human-readable reportzuvo/reports/benchmark-NNN-[run_id].json— machine-readable JSON per benchmark-output-schema.md
The JSON file must conform to the schema version "2.0" defined in ../../shared/includes/benchmark-output-schema.md.
JSON Output Structure
{
"version": "2.0",
"skill": "benchmark",
"run_id": "<run_id>",
"timestamp": "<ISO-8601>",
"project": "<project path>",
"mode": "default",
"task_source": "user",
"task_hash": "<64-char SHA-256>",
"task_snapshot": "<first 30000 chars>",
"options": { "with_tests": false, "with_adversarial": false, "with_static_checks": false },
"providers_attempted": ["claude", "gemini", "codex-fast"],
"providers_succeeded": ["claude", "gemini", "codex-fast"],
"scored": ["claude", "gemini", "codex-fast"],
"leaderboard": [...],
"scorecards": {...},
"meta_judge_model": "claude-opus-5",
"judge_presentation_order": ["gemini", "claude", "codex-fast"],
"judge_input_truncated": false
}
Run Log
Run: <ISO-8601-Z> benchmark <project> - - <VERDICT> <providers_count>-providers <mode> <notes> <BRANCH> <SHA7> <INCLUDES> <TIER>
Retrospective (REQUIRED)
Follow the retrospective protocol from retrospective.md.
Gate check → structured questions → TSV emit → markdown append.
If gate check skips: print "RETRO: skipped (trivial session)" and proceed.
After printing this block, append the Run: line value (without the Run: prefix) to the log file path resolved per ../../shared/includes/run-logger.md.
VERDICT: PASS (all providers scored), PARTIAL (some failed), FAIL (all failed), UNSCORED (judge parse failed).
BENCHMARK COMPLETE Block
BENCHMARK COMPLETE
Run ID: [run_id]
Mode: [default|corpus]
Task: [hash prefix] — [first 60 chars or "corpus tasks"]
Providers: [N attempted / M succeeded / K scored]
Judge: [model name] (opposite-model rule)
Results saved:
zuvo/reports/benchmark-NNN-[run_id].md
zuvo/reports/benchmark-NNN-[run_id].json
Top provider: [name] (quality: [score], time: [Xs], cost: $[N])
Self-eval bias range: [min] to [max] (positive = overconfident)
Corpus Mode — Additional Phases (phases 5–8)
Corpus mode activates when --mode corpus is set. Phases 5–8 extend the default pipeline.
Architecture note: Phases 5–8 are orchestrator-managed — they are executed by the LLM running this skill, not by
scripts/benchmark.sh. The bash runner handles Round 1 dispatch and raw result collection only. The skill orchestrator manages multi-round flow: reading round files from--round-dir, interpolating prompts, re-dispatching providers, and scoring.
See Phase 5 (adversarial round), Phase 6 (Round 2 test writing), Phase 7 (test scoring), and Phase 8 (corpus leaderboard) defined in the corpus mode extension below.
Corpus mode quality formula (both rounds):
quality = round((code_composite + test_composite) * 2.5) → range 0–100
Default mode formula (code only):
quality = round(code_composite * 5) → range 0–100
Corpus Mode Extension
This section only runs when --mode corpus is set.
Phase 5: Adversarial Cross-Review (if --with-adversarial)
Each provider's Round 1 output is stored as tmp/round1_<provider>.txt.
For each provider that succeeded in Phase 1:
- Run
scripts/adversarial-review.sh --files tmp/round1_<provider>.txt --json - Have the adversarial reviewers (other providers) challenge the code
- Record the adversarial findings; provider may rewrite →
tmp/round1_fixed_<provider>.txt - Compute
adversarial_delta: re-score the post-adversarial output. Delta = fixed_score − original_score (typically negative: critiques lower scores). - CQ8: if adversarial fails for a provider, continue with unfixed
round1_<provider>.txtoutput
Phase 6: Round 3 — Write Tests (if --with-tests)
Round 1 produces code. Round 3 produces tests for that code. (Round 2 is adversarial review — optional, can run between rounds.)
For each provider that succeeded Round 1:
- Load
../../shared/includes/benchmark-corpus/task-tests.md - Interpolate
{{ROUND_1_CODE}}with the provider'sround1_<provider>.txtoutput - Dispatch the interpolated prompt back to the same provider
- Capture Round 3 response (test code) →
tmp/round3_<provider>.txt+TEST_EVAL_SUMMARYblock
Phase 7: Round 4 — Adversarial on Tests (if --with-test-adversarial)
Mirrors Phase 5 but targets test files from Round 3.
Each provider's Round 3 output is stored as tmp/round3_<provider>.txt.
For each provider that succeeded Round 3:
- Run
scripts/adversarial-review.sh --files tmp/round3_<provider>.txt --json - Other providers critique: missing branches, weak assertions, echo tests, tautological oracles
- Author may rewrite tests →
tmp/round3_fixed_<provider>.txt - CQ8: if adversarial fails for a provider, continue with unfixed
round3_<provider>.txt
test_adversarial_delta = meta-judge score of fixed tests − original tests (typically negative).
Phase 7b: Score Round 3 Responses
Use the same meta-judge model to score test quality on four dimensions (0–5 each):
test_completeness: coverage of happy paths, error paths, edge casestest_accuracy: assertions are meaningful; mocks are correctly set uptest_actionability: tests are runnable and follow project conventionstest_no_hallucinations: no invented test utilities or non-existent matchers
test_composite = test_completeness + test_accuracy + test_actionability + test_no_hallucinations
Phase 8: Corpus Leaderboard
Assemble final leaderboard using the both-rounds quality formula:
quality = round((code_composite + test_composite) * 2.5)
Include test_score, adversarial_delta, tests_pass in the leaderboard and scorecards.
Print extended leaderboard with all columns populated.