Exp Bench Eval

Evaluates AI agents' end-to-end capability to conduct real AI research experiments, including designing methodologies, implementing code, executing experiments, and drawing conclusions. Use when the user wants to benchmark on EXP-Bench, or asks about evaluating this task. Reports All·E✓.

qhjqhj00 dbbb68a 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/exp-bench-eval commit dbbb68a2d9

Frequently asked questions

npx skillmds add qhjqhj00/exp-bench-eval