# Innovatorbench Eval

> innovatorbench-eval

- Skill: `qhjqhj00/innovatorbench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/innovatorbench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/innovatorbench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/innovatorbench-eval

---


# innovatorbench-eval

> InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research — Wu et al. (2025) (arXiv:2510.27598, 2025)

## What this evaluates

Evaluates AI agents' ability to conduct end-to-end LLM research across six domains: data construction, filtering, augmentation, loss/reward design, and scaffold construction. It probes long-horizon decision making, algorithmic robustness, resource management, and iterative code generation in a simulated research environment.

## Datasets

- **InnovatorBench** — total 20; splits: test (20); repo https://github.com/GAIR-NLP/InnovatorBench

## Metrics

- `Best Score` **(primary)** — range: [0, 100]
  - Highest achieved score among 3 evaluations and the final evaluation. Scores are computed automatically based on code execution and correctness checks per task domain.
- `Final Score` — range: [0, 100]
  - Score from the agent's last submission.

## Input / output format

**Input**: Task description, relevant dataset, and a starter code repository provided in a clean working directory inside a Docker container (Ubuntu 22.04). Agents can optionally access internet/tools depending on the task.

**Output**: Runnable Python code/scripts (e.g., data processing pipelines, loss/reward functions, training/inference scaffolds) submitted iteratively to the environment for execution.

## Scoring recipe

```python
def compute_scores(submissions, gold):
    final_score = run_and_evaluate(submissions[-1], gold)
    eval_scores = [run_and_evaluate(sub, gold) for sub in submissions[:3]]
    best_score = max(eval_scores + [final_score])
    return {'final_score': final_score, 'best_score': best_score}
```

## Common pitfalls

- Algorithm-related tasks (Loss/Reward Design) are highly brittle; minor code errors cause catastrophic failures like gradient explosion, unlike robust data tasks.
- Providing ground truth hints can hurt performance on data tasks because models struggle with exact replication and small implementation mismatches.
- Agents frequently mismanage computational resources (e.g., GPU contention) or terminate long training runs prematurely due to impatience.

## Evidence (verbatim from paper)

> Final Score: last submission score; Best Score: highest achieved score among 3 evaluations and final evaluation.

## Citation

```bibtex
@misc{wu2025innovatorbench,
  title={InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research},
  author={Wu et al. (2025)},
  year={2025},
  note={arXiv:2510.27598}
}
```

- arXiv: 2510.27598

