Benchmark

Measure whether a harness component (a skill, agent, or rule) actually earns its keep — run the same task k times with and without it, in isolated git worktrees, and report pass@k (works once) and pass^k (works consistently). Opt-in; not wired into the default pipeline. Use to decide keep-vs-cut on evidence, or to fill the R5 benchmark for a /extract proposal.

vasuag09 Updated

File contents

vasuag09/harness-claude/tree/main/skills/benchmark commit c25a226442

Frequently asked questions

npx skillmds@latest add vasuag09/benchmark