Researchgym Eval

Evaluates the capability of LLM agents to conduct closed-loop scientific research by proposing hypotheses, executing experiments, and outperforming human baselines on repurposed real-world AI papers. It probes long-horizon planning, resource management, and autonomous experimentation under realistic tool constraints. Use when the user wants to benchmark on ResearchGym, or asks about evaluating this task. Reports improvement over baselines.

qhjqhj00 9949e81 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/researchgym-eval commit 9949e81610

Frequently asked questions

npx skillmds add qhjqhj00/researchgym-eval