Autoresearch Prep
Invoke as $autoresearch-prep.
Scaffolds a program.md research program for $autoresearch. Auto-detects benchmark commands, test commands, sandbox boundaries, and exclude patterns from the codebase, then asks the user only what it can't infer: metric command and direction, research directions, and budget.
When to Use
When the user wants to start an $autoresearch loop but doesn't have a program.md yet.
Preconditions
- A git repository with code to optimize
- The user has a measurable property in mind (benchmark, bundle size, test coverage, etc.)
Process
0. Resolve Output Path
- If
$ARGUMENTS is provided and non-empty, use it as the output path.
- Otherwise default to
./program.md.
- If the file already exists, warn the user and ask whether to overwrite or pick a new path.
1. Auto-Detect Codebase Signals
Scan the repository for:
Benchmark commands:
package.json scripts containing bench or benchmark
bench/, benchmarks/, or perf/ directories
Makefile targets: bench, benchmark, perf
- Language-specific:
cargo bench, go test -bench, pytest-benchmark, hyperfine usage
Test commands:
package.json test script
Makefile test target
- Language-specific:
pytest, cargo test, go test, jest, vitest, mocha
Sandbox globs (based on what directories exist):
src/**, lib/**, app/**, pkg/**, internal/**, crates/**
- Include only directories that actually exist in the repo
Exclude patterns:
- Test files:
**/*.test.*, **/*.spec.*, **/test/**, **/tests/**, **/__tests__/**
- Config:
*.config.*, .eslintrc*, tsconfig*, .prettierrc*
- Benchmark harness:
bench/**, benchmarks/**, perf/**
- Build output:
dist/**, build/**, target/**, out/**, .next/**
- Lock files:
package-lock.json, yarn.lock, pnpm-lock.yaml, Cargo.lock, go.sum
- VCS/tooling:
.git/**, node_modules/**, .autoresearch/**
2. Light Interview
Present the detected signals as a summary, then ask:
- Metric command + direction (always required): "What shell command prints your target metric as a single number? Is higher or lower better?"
- Research directions (always required): "What optimization directions should the agent explore? List 2-5 ideas."
- Budget (offer default): "How many iterations? (default: 20)"
If benchmark commands were detected, suggest them as starting points for the metric command.
If test commands were detected, note they'll be included as the test command with required: true.
3. Findings Validation
Show the complete proposed program.md content to the user. Let them confirm or adjust any section before writing.
4. Validate Metric Command
- Run the metric command once from the project root.
- Parse the last number from stdout.
- If the command fails or produces no parseable number, stop and ask the user to fix it.
- Record the parsed value as the baseline for the Context section.
5. Write program.md
Write the file in the exact format expected by $autoresearch:
# Research Program
## Metric
command: <user-provided command>
direction: <higher-is-better | lower-is-better>
## Budget
max_iterations: <user-provided or 20>
## Sandbox
files:
- <detected or user-adjusted globs>
exclude:
- <detected or user-adjusted patterns>
## Test Command
command: <detected test command>
required: true
## Research Directions
1. <direction one>
2. <direction two>
...
## Context
Baseline metric value: <measured value>
<any additional context from interview>
Omit the Test Command section if no test command was detected or provided.
6. Route
Recommended next command: $autoresearch <output-path>
Constraints
- Do not run the autoresearch loop — only scaffold
program.md.
- Do not modify any source files in the repository.
- Do not create the
.autoresearch/ directory.
- Do not create alignment pages.
- Do not follow the shared shipping contract — this skill writes one local file, not tracked repo changes.
1---2name: autoresearch-prep3description: Scaffolds a program.md research program for autoresearch by auto-detecting codebase signals and interviewing for missing details.4---56# Autoresearch Prep78Invoke as `$autoresearch-prep`.910Scaffolds a `program.md` research program for `$autoresearch`. Auto-detects benchmark commands, test commands, sandbox boundaries, and exclude patterns from the codebase, then asks the user only what it can't infer: metric command and direction, research directions, and budget.1112## When to Use1314When the user wants to start an `$autoresearch` loop but doesn't have a `program.md` yet.1516## Preconditions1718- A git repository with code to optimize19- The user has a measurable property in mind (benchmark, bundle size, test coverage, etc.)2021## Process2223### 0. Resolve Output Path24251. If `$ARGUMENTS` is provided and non-empty, use it as the output path.262. Otherwise default to `./program.md`.273. If the file already exists, warn the user and ask whether to overwrite or pick a new path.2829### 1. Auto-Detect Codebase Signals3031Scan the repository for:3233**Benchmark commands:**34- `package.json` scripts containing `bench` or `benchmark`35- `bench/`, `benchmarks/`, or `perf/` directories36- `Makefile` targets: `bench`, `benchmark`, `perf`37- Language-specific: `cargo bench`, `go test -bench`, `pytest-benchmark`, `hyperfine` usage3839**Test commands:**40- `package.json` `test` script41- `Makefile` `test` target42- Language-specific: `pytest`, `cargo test`, `go test`, `jest`, `vitest`, `mocha`4344**Sandbox globs** (based on what directories exist):45- `src/**`, `lib/**`, `app/**`, `pkg/**`, `internal/**`, `crates/**`46- Include only directories that actually exist in the repo4748**Exclude patterns:**49- Test files: `**/*.test.*`, `**/*.spec.*`, `**/test/**`, `**/tests/**`, `**/__tests__/**`50- Config: `*.config.*`, `.eslintrc*`, `tsconfig*`, `.prettierrc*`51- Benchmark harness: `bench/**`, `benchmarks/**`, `perf/**`52- Build output: `dist/**`, `build/**`, `target/**`, `out/**`, `.next/**`53- Lock files: `package-lock.json`, `yarn.lock`, `pnpm-lock.yaml`, `Cargo.lock`, `go.sum`54- VCS/tooling: `.git/**`, `node_modules/**`, `.autoresearch/**`5556### 2. Light Interview5758Present the detected signals as a summary, then ask:59601. **Metric command + direction** (always required): "What shell command prints your target metric as a single number? Is higher or lower better?"612. **Research directions** (always required): "What optimization directions should the agent explore? List 2-5 ideas."623. **Budget** (offer default): "How many iterations? (default: 20)"6364If benchmark commands were detected, suggest them as starting points for the metric command.65If test commands were detected, note they'll be included as the test command with `required: true`.6667### 3. Findings Validation6869Show the complete proposed `program.md` content to the user. Let them confirm or adjust any section before writing.7071### 4. Validate Metric Command72731. Run the metric command once from the project root.742. Parse the last number from stdout.753. If the command fails or produces no parseable number, stop and ask the user to fix it.764. Record the parsed value as the baseline for the Context section.7778### 5. Write program.md7980Write the file in the exact format expected by `$autoresearch`:8182```markdown83# Research Program8485## Metric86command: <user-provided command>87direction: <higher-is-better | lower-is-better>8889## Budget90max_iterations: <user-provided or 20>9192## Sandbox93files:94 - <detected or user-adjusted globs>95exclude:96 - <detected or user-adjusted patterns>9798## Test Command99command: <detected test command>100required: true101102## Research Directions1031. <direction one>1042. <direction two>105...106107## Context108Baseline metric value: <measured value>109<any additional context from interview>110```111112Omit the `Test Command` section if no test command was detected or provided.113114### 6. Route115116**Recommended next command:** `$autoresearch <output-path>`117118## Constraints119120- Do not run the autoresearch loop — only scaffold `program.md`.121- Do not modify any source files in the repository.122- Do not create the `.autoresearch/` directory.123- Do not create alignment pages.124- Do not follow the shared shipping contract — this skill writes one local file, not tracked repo changes.