Rebuild Benchmark Audit

Audit and construct verifier-graded "rebuild the program" benchmarks and RLVR reward environments — hunting duplicate reward, shortcut-passable tests, ambient implementations already installed in the task image, proxy/short-circuit assertions, wall-clock time bombs, unreachable tests with missing fixtures, and grader-only dependencies. Use this whenever the user mentions ProgramBench, SWE-style or agentic benchmark construction, cleanroom task images, gold/dummy test validation, behavioral test suites as reward, verifier design, reward hacking, test-suite deduplication, or asks why a model scores suspiciously high or low on a coding benchmark — even if they don't use the word "benchmark". Also use it when designing the same recipe for APIs, libraries, protocol daemons, or data pipelines.

daedalus 479b9e7 23.1 KB Updated

File contents

daedalus/skills/tree/main/skills/program-bench-vetted commit 479b9e711e

Frequently asked questions

npx skillmds@latest add daedalus/rebuild-benchmark-audit