autoresearch
Autoresearch is a closed-loop ML experimentation workflow:
- human writes
program.md
- agent edits
train.py
prepare.py stays fixed
- every run gets the same 300-second budget
- lower
val_bpb wins
- regressions get reverted
This skill should behave like a routing-first front door, not a giant tutorial. Pick the user's mode, enforce the immutable-harness rules, then hand them to the smallest useful script or reference.
When to use this skill
- Set up
karpathy/autoresearch on a real GPU machine
- Write or refine
program.md before a session
- Run a bounded overnight
train.py search loop
- Interpret
results.tsv after a session
- Adapt the workflow to tighter VRAM constraints without invalidating comparisons
- Explain the ML-specific boundary between
autoresearch and nearby eval tooling
Do not use this skill when
- The user wants to optimize a
SKILL.md, prompt, or repo-local workflow with frozen prompts/evals — use skill-autoresearch
- The user wants app-level tracing, dataset-backed LLM evals, feedback review, or observability — use LangSmith, Braintrust, Weave, Promptfoo, or similar tools
- The job does not involve a real training repo,
program.md, train.py, fixed runtime budget, and val_bpb keep/revert ratcheting
- The user is really asking for a paper survey, general benchmark scan, or literature review with no intention to run the training loop
Core boundary
| Concern |
autoresearch owns |
Route elsewhere |
| Mutable target |
train.py in a real training repo |
prompts, app configs, SKILL.md, product behavior |
| Fixed evaluator |
prepare.py, validation shard, TIME_BUDGET=300, chosen MAX_SEQ_LEN / EVAL_TOKENS for the session |
prompt/eval datasets, app scorecards, observability dashboards |
| Acceptance rule |
keep only lower val_bpb; revert ties/regressions |
human review queues, app-level release gates |
| Main artifacts |
program.md, results.tsv, kept/discarded commits |
prompt suites, traces, feedback datasets |
If that boundary does not fit, do not stretch this skill.
Required intake packet
Before acting, identify:
- Mode — setup,
program.md, run loop, results interpretation, or constrained hardware
- Repository state — cloned or not, dependencies installed or not
- Hardware state — GPU / VRAM / CUDA / MLX / Windows path
- Session state — first baseline, active loop, or completed run
- Constraint state — target VRAM ceiling, whether
prepare.py has already been frozen for this session
Instructions
Step 1: Pick exactly one operating mode
Choose the smallest mode that answers the request:
Setup readiness
- install
uv
- clone repo
- sync dependencies
- verify GPU/CUDA/uv with
scripts/check-hardware.sh
- run the first baseline experiment
program.md authoring
- write or refine the human research charter
- record current baseline
val_bpb
- prioritize hypotheses
- list what has already been tried
- freeze constraints before the loop starts
Bounded run loop
- confirm the evaluator is already fixed
- use
train.py as the only mutable search surface
- run the loop with keep/revert discipline
- log every experiment to
results.tsv
Results interpretation
- summarize best kept runs
- identify repeated failures or crash patterns
- extract what belongs in the next
program.md
- distinguish genuine gains from one-off anomalies
Constrained-hardware adaptation
- set
MAX_SEQ_LEN and EVAL_TOKENS before the session
- keep them unchanged once the session starts
- adjust model/search strategy instead of cheating the evaluator mid-run
- route to community forks when CUDA assumptions do not hold
Do not answer all five modes at once unless the user explicitly asked for a full end-to-end walkthrough.
Step 2: Re-state the immutable harness
Every mode must preserve these rules:
program.md is human-authored and read-only during a session
train.py is the main mutable search surface
prepare.py is read-only once the session starts
TIME_BUDGET=300 stays fixed
val_bpb is the main keep/revert metric
results.tsv is append-only
- dependency set in
pyproject.toml stays locked
If the user wants to change the evaluator, start a new comparison track, not the current session.
Step 3: Execute the chosen mode
Mode A — Setup readiness
Use this path when the repo is not yet runnable.
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/karpathy/autoresearch
cd autoresearch
uv sync
bash scripts/check-hardware.sh
uv run prepare.py
uv run train.py > run.log 2>&1
grep "^val_bpb:\|^peak_vram_mb:" run.log
Success condition: one baseline run completes and prints both val_bpb and peak_vram_mb.
Mode B — program.md authoring
Use this path when the loop exists but direction is weak.
Minimum sections:
- goal tied to lower
val_bpb
- current baseline
val_bpb
- directions to explore in priority order
- what has been tried already
- constraints:
TIME_BUDGET=300, no prepare.py mutation, no new packages, VRAM ceiling, one meaningful change per experiment
For fuller templates and update patterns, use references/program-md-guide.md.
Mode C — Bounded run loop
Use this path only after setup and program.md are ready.
Loop contract:
- read
program.md + current train.py
- form one hypothesis
- edit
train.py
- commit
- run one 300-second experiment
- extract
val_bpb
- keep if improved, otherwise
git reset HEAD~1
- append result to
results.tsv
Typical commands:
bash scripts/run-experiment.sh
bash scripts/run-loop.sh --max 20 --desc "session-1"
Do not encourage multi-change hero rewrites. Clean ablations matter more than flashy edits.
Mode D — Results interpretation
Use this path after a completed run or checkpoint.
Helpful commands:
bash scripts/show-results.sh --top 10
awk -F'\t' '$4=="keep"' results.tsv | sort -t$'\t' -k2 -n
awk -F'\t' '{print $4}' results.tsv | sort | uniq -c
Summarize only four things: best gains, repeated failures, what should move into What Has Been Tried, and the next narrow experiment family.
Mode E — Constrained-hardware adaptation
Use this path when VRAM, platform, or runtime constraints dominate.
Rules:
- choose
MAX_SEQ_LEN and EVAL_TOKENS before the session
- never change them mid-session
- lower model/search ambition before mutating the evaluator
- prefer route-outs to community forks for Apple Silicon / non-CUDA paths
For concrete values and troubleshooting, use references/hardware-config.md.
Step 4: Route out aggressively when the request is adjacent
Route out when:
- the user wants to optimize instructions, prompts, or repo-local skills →
skill-autoresearch
- the user wants app-level traces, feedback review, observability, or online/offline eval dashboards → LangSmith / Braintrust / Weave / Promptfoo
- the user wants general literature synthesis rather than a runnable ML loop → research or survey tooling
Step 5: Keep the heavy detail in support files
Use support files instead of re-explaining everything inline:
references/operating-modes-and-route-outs.md — fast routing table, minimal response shape, and handoff logic
references/architecture.md — immutability contract, file map, metric rationale
references/program-md-guide.md — templates and update rules
references/hardware-config.md — VRAM tables and platform troubleshooting
scripts/*.sh — runnable setup / loop / reporting helpers
Available scripts
Run from inside the autoresearch repository directory:
| Script |
Purpose |
Usage |
setup.sh |
One-time environment setup |
bash scripts/setup.sh [--seq-len 512] |
run-experiment.sh |
Single 5-minute experiment + metric extraction |
bash scripts/run-experiment.sh |
run-loop.sh |
Autonomous loop: run → keep/revert → repeat |
bash scripts/run-loop.sh [--max 20] |
show-results.sh |
Human-readable results.tsv report |
bash scripts/show-results.sh [--top 10] |
check-hardware.sh |
GPU/CUDA/uv readiness check (JSON output) |
bash scripts/check-hardware.sh |
References
Detailed documentation in references/:
| File |
Contents |
references/operating-modes-and-route-outs.md |
Mode picker, adjacency boundaries, and minimal output contract |
references/architecture.md |
System design, immutability contract, git ratcheting, metric rationale |
references/program-md-guide.md |
How to write and update effective program.md directives |
references/hardware-config.md |
VRAM settings by GPU, memory optimization, platform troubleshooting |
Examples
Example 1: First 40GB GPU session
Request: “Help me run Karpathy autoresearch on a 40GB GPU.”
Expected behavior:
- choose Setup readiness first
- verify hardware and dependencies
- run one baseline experiment
- route to
program.md authoring only after the baseline exists
Example 2: User wants to optimize a skill instead
Request: “Can autoresearch help me improve this SKILL.md with binary evals?”
Expected behavior:
- route out immediately to
skill-autoresearch
- explain that this skill is for real ML training search on
train.py
Best practices
- Start with the smallest mode that fits — setup, authoring, run loop, interpretation, or hardware adaptation
- Baseline before bravado — confirm one successful run before talking about overnight loops
- Freeze the evaluator before the session —
prepare.py, TIME_BUDGET, MAX_SEQ_LEN, and EVAL_TOKENS must stay comparable
- One meaningful experiment at a time — ablations beat mystery bundles
- Keep
results.tsv append-only — discarded runs are still evidence
- Push deep detail into references/scripts — the front door should classify and route, not duplicate every table
- Route adjacent jobs away early — prompt/app eval and
SKILL.md optimization are different lanes
References
1---2name: autoresearch3description: Run Karpathy-style autonomous ML search on a real training repo: choose the right mode (setup, program.md, bounded loop, results interpretation, or constrained-hardware adaptation), preserve the immutable prepare.py / 300-second / val_bpb contract, and route prompt/skill eval work away to LangSmith, Promptfoo, Braintrust, or skill-autoresearch.4---567891011121314151617# autoresearch1819Autoresearch is a **closed-loop ML experimentation workflow**:20- human writes `program.md`21- agent edits `train.py`22- `prepare.py` stays fixed23- every run gets the same 300-second budget24- lower `val_bpb` wins25- regressions get reverted2627This skill should behave like a **routing-first front door**, not a giant tutorial. Pick the user's mode, enforce the immutable-harness rules, then hand them to the smallest useful script or reference.2829## When to use this skill3031- Set up `karpathy/autoresearch` on a real GPU machine32- Write or refine `program.md` before a session33- Run a bounded overnight `train.py` search loop34- Interpret `results.tsv` after a session35- Adapt the workflow to tighter VRAM constraints without invalidating comparisons36- Explain the ML-specific boundary between `autoresearch` and nearby eval tooling3738## Do not use this skill when3940- The user wants to optimize a `SKILL.md`, prompt, or repo-local workflow with frozen prompts/evals — use `skill-autoresearch`41- The user wants app-level tracing, dataset-backed LLM evals, feedback review, or observability — use LangSmith, Braintrust, Weave, Promptfoo, or similar tools42- The job does not involve a real training repo, `program.md`, `train.py`, fixed runtime budget, and `val_bpb` keep/revert ratcheting43- The user is really asking for a paper survey, general benchmark scan, or literature review with no intention to run the training loop4445## Core boundary4647| Concern | `autoresearch` owns | Route elsewhere |48|---------|---------------------|-----------------|49| Mutable target | `train.py` in a real training repo | prompts, app configs, `SKILL.md`, product behavior |50| Fixed evaluator | `prepare.py`, validation shard, `TIME_BUDGET=300`, chosen `MAX_SEQ_LEN` / `EVAL_TOKENS` for the session | prompt/eval datasets, app scorecards, observability dashboards |51| Acceptance rule | keep only lower `val_bpb`; revert ties/regressions | human review queues, app-level release gates |52| Main artifacts | `program.md`, `results.tsv`, kept/discarded commits | prompt suites, traces, feedback datasets |5354If that boundary does not fit, do not stretch this skill.5556## Required intake packet5758Before acting, identify:591. **Mode** — setup, `program.md`, run loop, results interpretation, or constrained hardware602. **Repository state** — cloned or not, dependencies installed or not613. **Hardware state** — GPU / VRAM / CUDA / MLX / Windows path624. **Session state** — first baseline, active loop, or completed run635. **Constraint state** — target VRAM ceiling, whether `prepare.py` has already been frozen for this session6465## Instructions6667### Step 1: Pick exactly one operating mode6869Choose the smallest mode that answers the request:70711. **Setup readiness**72 - install `uv`73 - clone repo74 - sync dependencies75 - verify GPU/CUDA/uv with `scripts/check-hardware.sh`76 - run the first baseline experiment77782. **`program.md` authoring**79 - write or refine the human research charter80 - record current baseline `val_bpb`81 - prioritize hypotheses82 - list what has already been tried83 - freeze constraints before the loop starts84853. **Bounded run loop**86 - confirm the evaluator is already fixed87 - use `train.py` as the only mutable search surface88 - run the loop with keep/revert discipline89 - log every experiment to `results.tsv`90914. **Results interpretation**92 - summarize best kept runs93 - identify repeated failures or crash patterns94 - extract what belongs in the next `program.md`95 - distinguish genuine gains from one-off anomalies96975. **Constrained-hardware adaptation**98 - set `MAX_SEQ_LEN` and `EVAL_TOKENS` before the session99 - keep them unchanged once the session starts100 - adjust model/search strategy instead of cheating the evaluator mid-run101 - route to community forks when CUDA assumptions do not hold102103Do **not** answer all five modes at once unless the user explicitly asked for a full end-to-end walkthrough.104105### Step 2: Re-state the immutable harness106107Every mode must preserve these rules:108109- `program.md` is human-authored and read-only during a session110- `train.py` is the main mutable search surface111- `prepare.py` is read-only once the session starts112- `TIME_BUDGET=300` stays fixed113- `val_bpb` is the main keep/revert metric114- `results.tsv` is append-only115- dependency set in `pyproject.toml` stays locked116117If the user wants to change the evaluator, start a **new comparison track**, not the current session.118119### Step 3: Execute the chosen mode120121#### Mode A — Setup readiness122123Use this path when the repo is not yet runnable.124125```bash126curl -LsSf https://astral.sh/uv/install.sh | sh127git clone https://github.com/karpathy/autoresearch128cd autoresearch129uv sync130bash scripts/check-hardware.sh131uv run prepare.py132uv run train.py > run.log 2>&1133grep "^val_bpb:\|^peak_vram_mb:" run.log134```135136Success condition: one baseline run completes and prints both `val_bpb` and `peak_vram_mb`.137138#### Mode B — `program.md` authoring139140Use this path when the loop exists but direction is weak.141142Minimum sections:143- goal tied to lower `val_bpb`144- current baseline `val_bpb`145- directions to explore in priority order146- what has been tried already147- constraints: `TIME_BUDGET=300`, no `prepare.py` mutation, no new packages, VRAM ceiling, one meaningful change per experiment148149For fuller templates and update patterns, use `references/program-md-guide.md`.150151#### Mode C — Bounded run loop152153Use this path only after setup and `program.md` are ready.154155Loop contract:1561. read `program.md` + current `train.py`1572. form one hypothesis1583. edit `train.py`1594. commit1605. run one 300-second experiment1616. extract `val_bpb`1627. keep if improved, otherwise `git reset HEAD~1`1638. append result to `results.tsv`164165Typical commands:166167```bash168bash scripts/run-experiment.sh169bash scripts/run-loop.sh --max 20 --desc "session-1"170```171172Do not encourage multi-change hero rewrites. Clean ablations matter more than flashy edits.173174#### Mode D — Results interpretation175176Use this path after a completed run or checkpoint.177178Helpful commands:179180```bash181bash scripts/show-results.sh --top 10182awk -F'\t' '$4=="keep"' results.tsv | sort -t$'\t' -k2 -n183awk -F'\t' '{print $4}' results.tsv | sort | uniq -c184```185186Summarize only four things: best gains, repeated failures, what should move into `What Has Been Tried`, and the next narrow experiment family.187188#### Mode E — Constrained-hardware adaptation189190Use this path when VRAM, platform, or runtime constraints dominate.191192Rules:193- choose `MAX_SEQ_LEN` and `EVAL_TOKENS` **before** the session194- never change them mid-session195- lower model/search ambition before mutating the evaluator196- prefer route-outs to community forks for Apple Silicon / non-CUDA paths197198For concrete values and troubleshooting, use `references/hardware-config.md`.199200### Step 4: Route out aggressively when the request is adjacent201202Route out when:203- the user wants to optimize instructions, prompts, or repo-local skills → `skill-autoresearch`204- the user wants app-level traces, feedback review, observability, or online/offline eval dashboards → LangSmith / Braintrust / Weave / Promptfoo205- the user wants general literature synthesis rather than a runnable ML loop → research or survey tooling206207### Step 5: Keep the heavy detail in support files208209Use support files instead of re-explaining everything inline:210- `references/operating-modes-and-route-outs.md` — fast routing table, minimal response shape, and handoff logic211- `references/architecture.md` — immutability contract, file map, metric rationale212- `references/program-md-guide.md` — templates and update rules213- `references/hardware-config.md` — VRAM tables and platform troubleshooting214- `scripts/*.sh` — runnable setup / loop / reporting helpers215216## Available scripts217218Run from inside the autoresearch repository directory:219220| Script | Purpose | Usage |221|--------|---------|-------|222| `setup.sh` | One-time environment setup | `bash scripts/setup.sh [--seq-len 512]` |223| `run-experiment.sh` | Single 5-minute experiment + metric extraction | `bash scripts/run-experiment.sh` |224| `run-loop.sh` | Autonomous loop: run → keep/revert → repeat | `bash scripts/run-loop.sh [--max 20]` |225| `show-results.sh` | Human-readable `results.tsv` report | `bash scripts/show-results.sh [--top 10]` |226| `check-hardware.sh` | GPU/CUDA/uv readiness check (JSON output) | `bash scripts/check-hardware.sh` |227228## References229230Detailed documentation in `references/`:231232| File | Contents |233|------|----------|234| `references/operating-modes-and-route-outs.md` | Mode picker, adjacency boundaries, and minimal output contract |235| `references/architecture.md` | System design, immutability contract, git ratcheting, metric rationale |236| `references/program-md-guide.md` | How to write and update effective `program.md` directives |237| `references/hardware-config.md` | VRAM settings by GPU, memory optimization, platform troubleshooting |238239## Examples240241### Example 1: First 40GB GPU session242243Request: “Help me run Karpathy autoresearch on a 40GB GPU.”244245Expected behavior:246- choose **Setup readiness** first247- verify hardware and dependencies248- run one baseline experiment249- route to `program.md` authoring only after the baseline exists250251### Example 2: User wants to optimize a skill instead252253Request: “Can autoresearch help me improve this `SKILL.md` with binary evals?”254255Expected behavior:256- route out immediately to `skill-autoresearch`257- explain that this skill is for real ML training search on `train.py`258259## Best practices2602611. **Start with the smallest mode that fits** — setup, authoring, run loop, interpretation, or hardware adaptation2622. **Baseline before bravado** — confirm one successful run before talking about overnight loops2633. **Freeze the evaluator before the session** — `prepare.py`, `TIME_BUDGET`, `MAX_SEQ_LEN`, and `EVAL_TOKENS` must stay comparable2644. **One meaningful experiment at a time** — ablations beat mystery bundles2655. **Keep `results.tsv` append-only** — discarded runs are still evidence2666. **Push deep detail into references/scripts** — the front door should classify and route, not duplicate every table2677. **Route adjacent jobs away early** — prompt/app eval and `SKILL.md` optimization are different lanes268269## References270271- [GitHub — karpathy/autoresearch](https://github.com/karpathy/autoresearch)272- [Karpathy — A Recipe for Training Neural Networks](https://karpathy.github.io/2019/04/25/recipe/)273- [MLflow Tracking](https://mlflow.org/docs/latest/ml/tracking/)274- [Weights & Biases Tracking](https://docs.wandb.ai/guides/track/)275- [MIT License](https://github.com/karpathy/autoresearch/blob/master/LICENSE)