Surfaces metaharness-darwin bench <create|verify> — the supporting verb
for harness-evolve --bench. Use when you want evolution scored against a
fixed corpus (independent of npm test) so champion fitness is comparable
across commits or across forks of the same harness.
When to use
- Setting up a new evolution pipeline for a repo whose
npm testis flaky, slow, or undersized — scaffold a deterministic bench suite once, then evolve against it repeatedly. - CI:
bench verifythe checked-in suite on every PR that touches it (cheap; ~5s). - Forking a harness to a new domain: copy and edit the suite to retarget the evaluation without losing comparability to the parent.
Algorithm
Implementation: scripts/bench.mjs.
--op create
- Resolve
--repopath; reject if missing. - Shell to
metaharness-darwin bench create <repo> [--out <suite.json>]. - Default output path:
<repo>/.metaharness/bench/suite.json(chosen by upstream). - Suite shape (per upstream): array of
{ input, expectedOutput, weight }tasks derived from existing test cases.
--op verify
- Resolve
--suitepath; reject if missing. - Shell to
metaharness-darwin bench verify <suite.json>. - Exit 1 if any task malformed (upstream's signal).
Output shape
{
"success": true,
"data": {
"op": "verify",
"taskCount": 42,
"wellFormed": true,
"durationMs": 870
}
}
Exit codes
| Code | Meaning |
|---|---|
| 0 | OK (or degraded — Darwin absent) |
| 1 | --op verify and suite malformed |
| 2 | Config error or upstream invocation failure |
Graceful degradation
When @metaharness/darwin is absent, emits the standard {degraded: true, reason: 'metaharness-darwin-not-available'} payload and exits 0.
Source: ruvnet/ruflo → plugins/ruflo-metaharness/skills/harness-bench/SKILL.md