Autoresearch
Invoke as $autoresearch.
Autonomous experiment loop inspired by Karpathy's autoresearch pattern. Reads a program.md research program, proposes one hypothesis per iteration, implements it within a declared sandbox, measures a single metric, and keeps only improvements. Git branches isolate experiments; a TSV log tracks all results.
When to Use
When the user wants to autonomously optimize a measurable codebase property: benchmark throughput, bundle size, test coverage, load time, memory usage, build speed, or any metric expressible as a shell command that prints a number.
Preconditions
- A
program.mdfile exists (default:./program.md, or path from$ARGUMENTS). If missing, recommend$autoresearch-prepto scaffold one. - Git working tree is clean (no uncommitted changes)
- The metric command runs successfully and prints a parseable number
program.md Format
The user creates this file. Required sections:
# Research Program
## Metric
command: <shell command that prints a single number to stdout>
direction: higher-is-better | lower-is-better
## Budget
max_iterations: <number>
## Sandbox
files:
- <glob patterns of files the agent may modify>
exclude:
- <glob patterns to never touch>
## Research Directions
1. <direction one>
2. <direction two>
Optional sections:
## Test Command
command: <shell command that must exit 0 before measuring>
required: true
## Budget
metric_timeout_seconds: <default 300>
## Context
<free-form notes about the codebase, prior work, constraints>
Process
0. Validate Preconditions
- Resolve
program.mdpath from$ARGUMENTS(default:./program.md). If the file does not exist, stop and recommend$autoresearch-prep. - Parse and validate required fields:
Metric.command,Metric.direction,Budget.max_iterations,Sandbox.files. - Verify git working tree is clean (
git status --porcelainis empty). Stop if dirty. - Create
.autoresearch/directory if it does not exist. - Record the current branch as the base branch.
- Add
.autoresearch/to.gitignoreif not already present.
1. Establish Baseline
- Run the metric command from the project root.
- Parse the last number from stdout (strip non-numeric lines).
- Record as iteration 0 in
.autoresearch/results.tsv:iteration timestamp hypothesis metric_value delta delta_pct status branch commit notes 0 <ISO-8601> baseline <value> 0 0.00% baseline <base-branch> <HEAD-sha> initial measurement - Set
best_value = baseline_value.
2. Check Stop Conditions
Before each iteration, check in order:
- Stop file: if
.autoresearch/stopexists, log "stop file detected" and go to step 9. - Iteration limit: if current iteration >
max_iterations, go to step 9. - Re-read program.md: pick up any user edits to directions, budget, or context mid-run.
3. Propose Hypothesis
- Read: research directions, past results from
.autoresearch/results.tsv, current sandbox file contents, and context section. - Propose exactly one specific, testable hypothesis. Prefer untried directions. Avoid repeating failed approaches.
- The hypothesis must name: what to change, which file(s), and the expected effect on the metric.
- Log the hypothesis before implementing.
4. Create Experiment Branch
- Ensure base branch is checked out and clean.
- Create and checkout:
autoresearch/iter-{N}-{slug}where{slug}is a 2-4 word kebab-case summary of the hypothesis.
5. Implement the Change
- Modify only files matching the sandbox glob patterns.
- Never modify files matching exclude patterns.
- Never modify: the metric command/script, benchmark suite, test harness, evaluation code, or anything outside the sandbox.
- Keep changes minimal and focused on the single hypothesis.
- Commit on the experiment branch with message:
autoresearch: iter {N} — {hypothesis summary}.
6. Validate Build/Tests
- If a test command is configured and
required: true:- Run the test command.
- If it fails: log
status: test-failed, checkout base branch, delete experiment branch, record in results.tsv, continue to next iteration.
- If no test command is configured, skip this step.
7. Measure
- Run the metric command with the configured timeout (default 300s).
- Parse the last number from stdout.
- If the command fails or times out: log
status: measure-failed, checkout base branch, delete experiment branch, record in results.tsv, continue to next iteration.
8. Evaluate (Ratchet)
Compare measured value to best_value using the configured direction:
If improved:
- Checkout base branch.
- Fast-forward merge the experiment branch:
git merge --ff-only autoresearch/iter-{N}-{slug}. - Update
best_value. - Record
status: keptwith the delta and delta percentage in results.tsv. - Delete the experiment branch (it's merged).
If not improved (or equal):
- Checkout base branch.
- Record
status: revertedwith the delta in results.tsv. - Delete the experiment branch.
Always ensure the working tree is clean before proceeding to the next iteration.
9. Report Final Status
When the loop ends (stop file, iteration limit, or all directions exhausted):
- Write
.autoresearch/summary.md:- Baseline value → final best value and total improvement (absolute + percentage).
- Per-iteration table (from results.tsv).
- Directions attempted vs. untried.
- Top 3 most impactful kept changes.
- Print a summary table to the terminal.
- Do not push to remote — all work stays local for user review.
Output
.autoresearch/results.tsv— append-only experiment log.autoresearch/summary.md— written at loop end
Constraints
- Sandbox is absolute: never modify files outside declared sandbox globs.
- Never touch eval: never modify the metric command, benchmark suite, or test harness.
- One hypothesis per iteration: no bundling multiple changes.
- Clean state between iterations: never leave dirty working tree.
- No remote push: all work stays local.
- No GitHub Actions: do not create CI workflows.
- Autonomous: no user approval per iteration. User steers by editing
program.mdor creating.autoresearch/stop.
Failure Recovery
If the agent finds itself on an experiment branch with uncommitted changes between iterations:
git checkout -- .to discard changes.git checkout <base-branch>.- Delete the orphaned experiment branch.
- Record
status: crashed-recoveredin results.tsv. - Continue from the next iteration.
Shipping
This skill does not follow the shared shipping contract. All work stays local and unpushed. The user reviews results and decides what to keep and push.
Next work: none — loop is self-contained
Recommended next command: git log --oneline to review kept changes, then push when satisfied