Mle Bench Lite Eval

Evaluates an agent's ability to iteratively refine and improve runnable solutions for competition-style ML tasks over a long horizon. It probes sustained experiment improvement and competitive performance rather than just initial submission validity. Use when the user wants to benchmark on MLE-Bench Lite, or asks about evaluating this task. Reports Any Medal%.

qhjqhj00 1a5b03f 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mle-bench-lite-eval commit 1a5b03fc81

Frequently asked questions

npx skillmds add qhjqhj00/mle-bench-lite-eval