# Scientific LLM Benchmarks

> A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery.

- Skill: `akillness/scientific-llm-benchmarks` (Agent Skill, multi-file: 6 files)
- Install (CLI): `npx skillmds@latest add akillness/scientific-llm-benchmarks`
- Raw SKILL.md: https://api.skillmd.com/api/skills/akillness/scientific-llm-benchmarks/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Product & Planning
- Author: akillness (https://skillmd.com/u/akillness)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/akillness/scientific-llm-benchmarks

---


# Scientific LLM Benchmarks

This skill provides references to benchmarks used for evaluating large language models on scientific reasoning and discovery. The data comes from the Awesome-Scientific-LLM-Benchmarks repository.

## Contents
The complete benchmark list is stored locally within this skill:
- **References List:** `references/benchmarks.md`
- **Data (YAML format):** `data/benchmarks.yaml` 

## Benchmark Domains Covered
- **General / Multi-domain Science:** Cross-disciplinary STEM reasoning benchmarks.
- **Mathematics:** Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.
- **Physics and Astronomy:** Physics olympiad, graduate physics, computational physics, and astronomy.
- **Chemistry:** Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.
- **Materials Science:** Crystals, materials property prediction, and materials-science knowledge.
- **Biology and Life Sciences:** Genomics, proteins, bioinformatics agents, protocols, and research biology.
- **Agentic Science and AI Research:** LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.

## Helper Scripts
Also included are python scripts inside `scripts/`:
- `generate_readme.py`: Regenerates the markdown tables and list from `data/benchmarks.yaml`.
- `fetch_examples.py`: Fetches real sample rows from HuggingFace dataset URLs specified in the dataset metadata.
