Airs Bench Eval

Evaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.

qhjqhj00 e4a5a88 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/airs-bench-eval commit e4a5a88d95

Frequently asked questions

npx skillmds add qhjqhj00/airs-bench-eval