Paperbench Eval

Evaluates an autonomous agent's ability to replicate top-tier conference ML papers from scratch. It probes long-horizon engineering capabilities by measuring performance across 20 diverse tasks under a strict 24-hour time and compute budget. Use when the user wants to benchmark on PaperBench, or asks about evaluating this task. Reports Average Score.

qhjqhj00 32628d3 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/paperbench-eval commit 32628d3506

Frequently asked questions

npx skillmds add qhjqhj00/paperbench-eval