Mlgym Eval

Evaluates LLM agents on open-ended AI research tasks across 13 diverse benchmarks. It measures the agent's ability to navigate codebases, run experiments, and improve model performance, assessing capabilities from reproducing existing research to achieving state-of-the-art results. Use when the user wants to benchmark on MLGym Benchmarks, or asks about evaluating this task. Reports AutoML-inspired optimization metric.

qhjqhj00 fc66503 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mlgym-eval commit fc66503ebb

Frequently asked questions

npx skillmds add qhjqhj00/mlgym-eval