Claude Autoresearch -- Autonomous Goal-directed Iteration
Inspired by Karpathy's autoresearch. Applies constraint-driven autonomous iteration to ANY work -- not just ML research.
Core idea: You are an autonomous agent. Modify -> Verify -> Keep/Discard -> Repeat.
Subcommands
| Subcommand |
Purpose |
/autoresearch |
Run the autonomous loop (default) |
/autoresearch:plan |
Interactive wizard to build Scope, Metric, Direction & Verify from a Goal |
/autoresearch:security |
Autonomous security audit: STRIDE threat model + OWASP Top 10 + red-team (4 adversarial personas) |
/autoresearch:ship |
Universal shipping workflow: ship code, content, marketing, sales, research, or anything |
/autoresearch:debug |
Autonomous bug-hunting loop: scientific method + iterative investigation until codebase is clean |
/autoresearch:fix |
Autonomous fix loop: iteratively repair errors (tests, types, lint, build) until zero remain |
/autoresearch:security -- Autonomous Security Audit (v1.0.3)
Runs a comprehensive security audit using the autoresearch loop pattern. Generates a full STRIDE threat model, maps attack surfaces, then iteratively tests each vulnerability vector -- logging findings with severity, OWASP category, and code evidence.
Load: references/security-workflow.md for full protocol.
What it does:
- Codebase Reconnaissance -- scans tech stack, dependencies, configs, API routes
- Asset Identification -- catalogs data stores, auth systems, external services, user inputs
- Trust Boundary Mapping -- browser<->server, public<->authenticated, user<->admin, CI/CD<->prod
- STRIDE Threat Model -- Spoofing, Tampering, Repudiation, Info Disclosure, DoS, Elevation of Privilege
- Attack Surface Map -- entry points, data flows, abuse paths
- Autonomous Loop -- iteratively tests each vector, validates with code evidence, logs findings
- Final Report -- severity-ranked findings with mitigations, coverage matrix, iteration log
Key behaviors:
- Follows red-team adversarial mindset (Security Adversary, Supply Chain, Insider Threat, Infra Attacker)
- Every finding requires code evidence (file:line + attack scenario) -- no theoretical fluff
- Tracks OWASP Top 10 + STRIDE coverage, prints coverage summary every 5 iterations
- Composite metric:
(owasp_tested/10)*50 + (stride_tested/6)*30 + min(findings, 20) -- higher is better
- Creates
security/{YYMMDD}-{HHMM}-{audit-slug}/ folder with structured reports:
overview.md, threat-model.md, attack-surface-map.md, findings.md, owasp-coverage.md, dependency-audit.md, recommendations.md, security-audit-results.tsv
Flags:
| Flag |
Purpose |
--diff |
Delta mode -- only audit files changed since last audit |
--fix |
After audit, auto-fix confirmed Critical/High findings using autoresearch loop |
--fail-on {severity} |
Exit non-zero if findings meet threshold (for CI/CD gating) |
Usage:
# Unlimited -- keep finding vulnerabilities until interrupted
/autoresearch:security
# Bounded -- exactly 10 security sweep iterations
/autoresearch:security
Iterations: 10
# With focused scope
/autoresearch:security
Scope: src/api/**/*.ts, src/middleware/**/*.ts
Focus: authentication and authorization flows
# Delta mode -- only audit changed files since last audit
/autoresearch:security --diff
# Auto-fix confirmed Critical/High findings after audit
/autoresearch:security --fix
Iterations: 15
# CI/CD gate -- fail pipeline if any Critical findings
/autoresearch:security --fail-on critical
Iterations: 10
# Combined -- delta audit + fix + gate
/autoresearch:security --diff --fix --fail-on critical
Iterations: 15
Inspired by:
- Strix -- AI-powered security testing with proof-of-concept validation
- OWASP Top 10 (2021) -- industry-standard vulnerability taxonomy
- STRIDE -- Microsoft's threat modeling framework
/autoresearch:ship -- Universal Shipping Workflow (v1.1.0)
Ship anything -- code, content, marketing, sales, research, or design -- through a structured 8-phase workflow that applies autoresearch loop principles to the last mile.
Load: references/ship-workflow.md for full protocol.
What it does:
- Identify -- auto-detect what you're shipping (code PR, deployment, blog post, email campaign, sales deck, research paper, design assets)
- Inventory -- assess current state and readiness gaps
- Checklist -- generate domain-specific pre-ship gates (all mechanically verifiable)
- Prepare -- autoresearch loop to fix failing checklist items until 100% pass
- Dry-run -- simulate the ship action without side effects
- Ship -- execute the actual delivery (merge, deploy, publish, send)
- Verify -- post-ship health check confirms it landed
- Log -- record shipment to
ship-log.tsv for traceability
Supported shipment types:
| Type |
Example Ship Actions |
code-pr |
gh pr create with full description |
code-release |
Git tag + GitHub release |
deployment |
CI/CD trigger, kubectl apply, push to deploy branch |
content |
Publish via CMS, commit to content branch |
marketing-email |
Send via ESP (SendGrid, Mailchimp) |
marketing-campaign |
Activate ads, launch landing page |
sales |
Send proposal, share deck |
research |
Upload to repository, submit paper |
design |
Export assets, share with stakeholders |
Flags:
| Flag |
Purpose |
--dry-run |
Validate everything but don't actually ship (stop at Phase 5) |
--auto |
Auto-approve dry-run gate if no errors |
--force |
Skip non-critical checklist items (blockers still enforced) |
--rollback |
Undo the last ship action (if reversible) |
--monitor N |
Post-ship monitoring for N minutes |
--type <type> |
Override auto-detection with explicit shipment type |
--checklist-only |
Only generate and evaluate checklist (stop at Phase 3) |
Usage:
# Auto-detect and ship (interactive)
/autoresearch:ship
# Ship code PR with auto-approve
/autoresearch:ship --auto
# Dry-run a deployment before going live
/autoresearch:ship --type deployment --dry-run
# Ship with post-deployment monitoring
/autoresearch:ship --monitor 10
# Prepare iteratively then ship
/autoresearch:ship
Iterations: 5
# Just check if something is ready to ship
/autoresearch:ship --checklist-only
# Ship a blog post
/autoresearch:ship
Target: content/blog/my-new-post.md
Type: content
# Ship a sales deck
/autoresearch:ship --type sales
Target: decks/q1-proposal.pdf
# Rollback a bad deployment
/autoresearch:ship --rollback
Composite metric (for bounded loops):
ship_score = (checklist_passing / checklist_total) * 80
+ (dry_run_passed ? 15 : 0)
+ (no_blockers ? 5 : 0)
Score of 100 = fully ready. Below 80 = not shippable.
Output directory: Creates ship/{YYMMDD}-{HHMM}-{ship-slug}/ with checklist.md, ship-log.tsv, summary.md.
/autoresearch:plan -- Goal -> Configuration Wizard
Converts a plain-language goal into a validated, ready-to-execute autoresearch configuration.
Load: references/plan-workflow.md for full protocol.
Quick summary:
- Capture Goal -- ask what the user wants to improve (or accept inline text)
- Analyze Context -- scan codebase for tooling, test runners, build scripts
- Define Scope -- suggest file globs, validate they resolve to real files
- Define Metric -- suggest mechanical metrics, validate they output a number
- Define Direction -- higher or lower is better
- Define Verify -- construct the shell command, dry-run it, confirm it works
- Confirm & Launch -- present the complete config, offer to launch immediately
Critical gates:
- Metric MUST be mechanical (outputs a parseable number, not subjective)
- Verify command MUST pass a dry run on the current codebase before accepting
- Scope MUST resolve to >=1 file
Usage:
/autoresearch:plan
Goal: Make the API respond faster
/autoresearch:plan Increase test coverage to 95%
/autoresearch:plan Reduce bundle size below 200KB
After the wizard completes, the user gets a ready-to-paste /autoresearch invocation -- or can launch it directly.
When to Activate
- User invokes
/autoresearch or /ug:autoresearch -> run the loop
- User invokes
/autoresearch:plan -> run the planning wizard
- User invokes
/autoresearch:security -> run the security audit
- User says "help me set up autoresearch", "plan an autoresearch run" -> run the planning wizard
- User says "security audit", "threat model", "OWASP", "STRIDE", "find vulnerabilities", "red-team" -> run the security audit
- User invokes
/autoresearch:ship -> run the ship workflow
- User says "ship it", "deploy this", "publish this", "launch this", "get this out the door" -> run the ship workflow
- User invokes
/autoresearch:debug -> run the debug loop
- User says "find all bugs", "hunt bugs", "debug this", "why is this failing", "investigate" -> run the debug loop
- User invokes
/autoresearch:fix -> run the fix loop
- User says "fix all errors", "make tests pass", "fix the build", "clean up errors" -> run the fix loop
- User says "work autonomously", "iterate until done", "keep improving", "run overnight" -> run the loop
- Any task requiring repeated iteration cycles with measurable outcomes -> run the loop
Bounded Iterations
By default, autoresearch loops forever until manually interrupted. To run exactly N iterations, add Iterations: N to your inline config.
Unlimited (default):
/autoresearch
Goal: Increase test coverage to 90%
Bounded (N iterations):
/autoresearch
Goal: Increase test coverage to 90%
Iterations: 25
After N iterations Claude stops and prints a final summary with baseline -> current best, keeps/discards/crashes. If the goal is achieved before N iterations, Claude prints early completion and stops.
When to Use Bounded Iterations
| Scenario |
Recommendation |
| Run overnight, review in morning |
Unlimited (default) |
| Quick 30-min improvement session |
Iterations: 10 |
| Targeted fix with known scope |
Iterations: 5 |
| Exploratory -- see if approach works |
Iterations: 15 |
| CI/CD pipeline integration |
--iterations N flag (set N based on time budget) |
Setup Phase (Do Once)
If the user provides Goal, Scope, Metric, and Verify inline -> extract them and proceed to step 5.
If any critical field is missing -> use AskUserQuestion to collect them interactively:
Interactive Setup (when invoked without full config)
Scan the codebase first for smart defaults, then ask ALL questions in batched AskUserQuestion calls (max 4 per call). This gives users full clarity upfront.
Batch 1 -- Core config (4 questions in one call):
Use a SINGLE AskUserQuestion call with these 4 questions:
| # |
Header |
Question |
Options (smart defaults from codebase scan) |
| 1 |
Goal |
"What do you want to improve?" |
"Test coverage (higher)", "Bundle size (lower)", "Performance (faster)", "Code quality (fewer errors)" |
| 2 |
Scope |
"Which files can autoresearch modify?" |
Suggested globs from project structure (e.g. "src//*.ts", "content//*.md") |
| 3 |
Metric |
"What number tells you if it got better? (must be a command output, not subjective)" |
Detected options: "coverage % (higher)", "bundle size KB (lower)", "error count (lower)", "test pass count (higher)" |
| 4 |
Direction |
"Higher or lower is better?" |
"Higher is better", "Lower is better" |
Batch 2 -- Verify + Guard + Launch (3 questions in one call):
| # |
Header |
Question |
Options |
| 5 |
Verify |
"What command produces the metric? (I'll dry-run it to confirm)" |
Suggested commands from detected tooling |
| 6 |
Guard |
"Any command that must ALWAYS pass? (prevents regressions)" |
"npm test", "tsc --noEmit", "npm run build", "Skip -- no guard" |
| 7 |
Launch |
"Ready to go?" |
"Launch (unlimited)", "Launch with iteration limit", "Edit config", "Cancel" |
After Batch 2: Dry-run the verify command. If it fails, ask user to fix or choose a different command. If it passes, proceed with launch choice.
IMPORTANT: Always batch questions -- never ask one at a time. Users should see all config choices together for full context.
Setup Steps (after config is complete)
- Read all in-scope files for full context before any modification
- Define the goal -- extracted from user input or inline config
- Define scope constraints -- validated file globs
- Define guard (optional) -- regression prevention command
- Create a results log -- Track every iteration (see
references/results-logging.md)
- Establish baseline -- Run verification on current state AND guard (if set). Record as iteration #0
- Confirm and go -- Show user the setup, get confirmation, then BEGIN THE LOOP
The Loop
Read references/autonomous-loop-protocol.md for full protocol details.
LOOP (FOREVER or N times):
1. Review: Read current state + git history + results log
2. Ideate: Pick next change based on goal, past results, what hasn't been tried
3. Modify: Make ONE focused change to in-scope files
4. Commit: Git commit the change (before verification)
5. Verify: Run the mechanical metric (tests, build, benchmark, etc.)
6. Guard: If guard is set, run the guard command
7. Decide:
- IMPROVED + guard passed (or no guard) -> Keep commit, log "keep", advance
- IMPROVED + guard FAILED -> Revert, then try to rework the optimization
(max 2 attempts) so it improves the metric WITHOUT breaking the guard.
Never modify guard/test files -- adapt the implementation instead.
If still failing -> log "discard (guard failed)" and move on
- SAME/WORSE -> Git revert, log "discard"
- CRASHED -> Try to fix (max 3 attempts), else log "crash" and move on
8. Log: Record result in results log
9. Repeat: Go to step 1.
- If unbounded: NEVER STOP. NEVER ASK "should I continue?"
- If bounded (N): Stop after N iterations, print final summary
Critical Rules
- Loop until done -- Unbounded: loop until interrupted. Bounded: loop N times then summarize.
- Read before write -- Always understand full context before modifying
- One change per iteration -- Atomic changes. If it breaks, you know exactly why
- Mechanical verification only -- No subjective "looks good". Use metrics
- Automatic rollback -- Failed changes revert instantly. No debates
- Simplicity wins -- Equal results + less code = KEEP. Tiny improvement + ugly complexity = DISCARD
- Git is memory -- Every kept change committed. Agent reads history to learn patterns
- When stuck, think harder -- Re-read files, re-read goal, combine near-misses, try radical changes. Don't ask for help unless truly blocked by missing access/permissions
Principles Reference
See references/core-principles.md for the 7 generalizable principles from autoresearch.
Adapting to Different Domains
| Domain |
Metric |
Scope |
Verify Command |
Guard |
| Backend code |
Tests pass + coverage % |
src/**/*.ts |
npm test |
-- |
| Frontend UI |
Lighthouse score |
src/components/** |
npx lighthouse |
npm test |
| ML training |
val_bpb / loss |
train.py |
uv run train.py |
-- |
| Blog/content |
Word count + readability |
content/*.md |
Custom script |
-- |
| Performance |
Benchmark time (ms) |
Target files |
npm run bench |
npm test |
| Refactoring |
Tests pass + LOC reduced |
Target module |
npm test && wc -l |
npm run typecheck |
| Security |
OWASP + STRIDE coverage + findings |
API/auth/middleware |
/autoresearch:security |
-- |
| Shipping |
Checklist pass rate (%) |
Any artifact |
/autoresearch:ship |
Domain-specific |
| Debugging |
Bugs found + coverage |
Target files |
/autoresearch:debug |
-- |
| Fixing |
Error count (lower) |
Target files |
/autoresearch:fix |
npm test |
Adapt the loop to your domain. The PRINCIPLES are universal; the METRICS are domain-specific.
1---2name: autoresearch3description: Claude Autoresearch -- Autonomous Goal-directed Iteration4---56# Claude Autoresearch -- Autonomous Goal-directed Iteration78Inspired by [Karpathy's autoresearch](https://github.com/karpathy/autoresearch). Applies constraint-driven autonomous iteration to ANY work -- not just ML research.910**Core idea:** You are an autonomous agent. Modify -> Verify -> Keep/Discard -> Repeat.1112## Subcommands1314| Subcommand | Purpose |15|------------|---------|16| `/autoresearch` | Run the autonomous loop (default) |17| `/autoresearch:plan` | Interactive wizard to build Scope, Metric, Direction & Verify from a Goal |18| `/autoresearch:security` | Autonomous security audit: STRIDE threat model + OWASP Top 10 + red-team (4 adversarial personas) |19| `/autoresearch:ship` | Universal shipping workflow: ship code, content, marketing, sales, research, or anything |20| `/autoresearch:debug` | Autonomous bug-hunting loop: scientific method + iterative investigation until codebase is clean |21| `/autoresearch:fix` | Autonomous fix loop: iteratively repair errors (tests, types, lint, build) until zero remain |2223### /autoresearch:security -- Autonomous Security Audit (v1.0.3)2425Runs a comprehensive security audit using the autoresearch loop pattern. Generates a full STRIDE threat model, maps attack surfaces, then iteratively tests each vulnerability vector -- logging findings with severity, OWASP category, and code evidence.2627Load: `references/security-workflow.md` for full protocol.2829**What it does:**30311. **Codebase Reconnaissance** -- scans tech stack, dependencies, configs, API routes322. **Asset Identification** -- catalogs data stores, auth systems, external services, user inputs333. **Trust Boundary Mapping** -- browser<->server, public<->authenticated, user<->admin, CI/CD<->prod344. **STRIDE Threat Model** -- Spoofing, Tampering, Repudiation, Info Disclosure, DoS, Elevation of Privilege355. **Attack Surface Map** -- entry points, data flows, abuse paths366. **Autonomous Loop** -- iteratively tests each vector, validates with code evidence, logs findings377. **Final Report** -- severity-ranked findings with mitigations, coverage matrix, iteration log3839**Key behaviors:**40- Follows red-team adversarial mindset (Security Adversary, Supply Chain, Insider Threat, Infra Attacker)41- Every finding requires **code evidence** (file:line + attack scenario) -- no theoretical fluff42- Tracks OWASP Top 10 + STRIDE coverage, prints coverage summary every 5 iterations43- Composite metric: `(owasp_tested/10)*50 + (stride_tested/6)*30 + min(findings, 20)` -- higher is better44- Creates `security/{YYMMDD}-{HHMM}-{audit-slug}/` folder with structured reports:45 `overview.md`, `threat-model.md`, `attack-surface-map.md`, `findings.md`, `owasp-coverage.md`, `dependency-audit.md`, `recommendations.md`, `security-audit-results.tsv`4647**Flags:**4849| Flag | Purpose |50|------|---------|51| `--diff` | Delta mode -- only audit files changed since last audit |52| `--fix` | After audit, auto-fix confirmed Critical/High findings using autoresearch loop |53| `--fail-on {severity}` | Exit non-zero if findings meet threshold (for CI/CD gating) |5455**Usage:**56```57# Unlimited -- keep finding vulnerabilities until interrupted58/autoresearch:security5960# Bounded -- exactly 10 security sweep iterations61/autoresearch:security62Iterations: 106364# With focused scope65/autoresearch:security66Scope: src/api/**/*.ts, src/middleware/**/*.ts67Focus: authentication and authorization flows6869# Delta mode -- only audit changed files since last audit70/autoresearch:security --diff7172# Auto-fix confirmed Critical/High findings after audit73/autoresearch:security --fix74Iterations: 157576# CI/CD gate -- fail pipeline if any Critical findings77/autoresearch:security --fail-on critical78Iterations: 107980# Combined -- delta audit + fix + gate81/autoresearch:security --diff --fix --fail-on critical82Iterations: 1583```8485**Inspired by:**86- [Strix](https://github.com/usestrix/strix) -- AI-powered security testing with proof-of-concept validation87- OWASP Top 10 (2021) -- industry-standard vulnerability taxonomy88- STRIDE -- Microsoft's threat modeling framework8990### /autoresearch:ship -- Universal Shipping Workflow (v1.1.0)9192Ship anything -- code, content, marketing, sales, research, or design -- through a structured 8-phase workflow that applies autoresearch loop principles to the last mile.9394Load: `references/ship-workflow.md` for full protocol.9596**What it does:**97981. **Identify** -- auto-detect what you're shipping (code PR, deployment, blog post, email campaign, sales deck, research paper, design assets)992. **Inventory** -- assess current state and readiness gaps1003. **Checklist** -- generate domain-specific pre-ship gates (all mechanically verifiable)1014. **Prepare** -- autoresearch loop to fix failing checklist items until 100% pass1025. **Dry-run** -- simulate the ship action without side effects1036. **Ship** -- execute the actual delivery (merge, deploy, publish, send)1047. **Verify** -- post-ship health check confirms it landed1058. **Log** -- record shipment to `ship-log.tsv` for traceability106107**Supported shipment types:**108109| Type | Example Ship Actions |110|------|---------------------|111| `code-pr` | `gh pr create` with full description |112| `code-release` | Git tag + GitHub release |113| `deployment` | CI/CD trigger, `kubectl apply`, push to deploy branch |114| `content` | Publish via CMS, commit to content branch |115| `marketing-email` | Send via ESP (SendGrid, Mailchimp) |116| `marketing-campaign` | Activate ads, launch landing page |117| `sales` | Send proposal, share deck |118| `research` | Upload to repository, submit paper |119| `design` | Export assets, share with stakeholders |120121**Flags:**122123| Flag | Purpose |124|------|---------|125| `--dry-run` | Validate everything but don't actually ship (stop at Phase 5) |126| `--auto` | Auto-approve dry-run gate if no errors |127| `--force` | Skip non-critical checklist items (blockers still enforced) |128| `--rollback` | Undo the last ship action (if reversible) |129| `--monitor N` | Post-ship monitoring for N minutes |130| `--type <type>` | Override auto-detection with explicit shipment type |131| `--checklist-only` | Only generate and evaluate checklist (stop at Phase 3) |132133**Usage:**134```135# Auto-detect and ship (interactive)136/autoresearch:ship137138# Ship code PR with auto-approve139/autoresearch:ship --auto140141# Dry-run a deployment before going live142/autoresearch:ship --type deployment --dry-run143144# Ship with post-deployment monitoring145/autoresearch:ship --monitor 10146147# Prepare iteratively then ship148/autoresearch:ship149Iterations: 5150151# Just check if something is ready to ship152/autoresearch:ship --checklist-only153154# Ship a blog post155/autoresearch:ship156Target: content/blog/my-new-post.md157Type: content158159# Ship a sales deck160/autoresearch:ship --type sales161Target: decks/q1-proposal.pdf162163# Rollback a bad deployment164/autoresearch:ship --rollback165```166167**Composite metric (for bounded loops):**168```169ship_score = (checklist_passing / checklist_total) * 80170 + (dry_run_passed ? 15 : 0)171 + (no_blockers ? 5 : 0)172```173Score of 100 = fully ready. Below 80 = not shippable.174175**Output directory:** Creates `ship/{YYMMDD}-{HHMM}-{ship-slug}/` with `checklist.md`, `ship-log.tsv`, `summary.md`.176177### /autoresearch:plan -- Goal -> Configuration Wizard178179Converts a plain-language goal into a validated, ready-to-execute autoresearch configuration.180181Load: `references/plan-workflow.md` for full protocol.182183**Quick summary:**1841851. **Capture Goal** -- ask what the user wants to improve (or accept inline text)1862. **Analyze Context** -- scan codebase for tooling, test runners, build scripts1873. **Define Scope** -- suggest file globs, validate they resolve to real files1884. **Define Metric** -- suggest mechanical metrics, validate they output a number1895. **Define Direction** -- higher or lower is better1906. **Define Verify** -- construct the shell command, **dry-run it**, confirm it works1917. **Confirm & Launch** -- present the complete config, offer to launch immediately192193**Critical gates:**194- Metric MUST be mechanical (outputs a parseable number, not subjective)195- Verify command MUST pass a dry run on the current codebase before accepting196- Scope MUST resolve to >=1 file197198**Usage:**199```200/autoresearch:plan201Goal: Make the API respond faster202203/autoresearch:plan Increase test coverage to 95%204205/autoresearch:plan Reduce bundle size below 200KB206```207208After the wizard completes, the user gets a ready-to-paste `/autoresearch` invocation -- or can launch it directly.209210## When to Activate211212- User invokes `/autoresearch` or `/ug:autoresearch` -> run the loop213- User invokes `/autoresearch:plan` -> run the planning wizard214- User invokes `/autoresearch:security` -> run the security audit215- User says "help me set up autoresearch", "plan an autoresearch run" -> run the planning wizard216- User says "security audit", "threat model", "OWASP", "STRIDE", "find vulnerabilities", "red-team" -> run the security audit217- User invokes `/autoresearch:ship` -> run the ship workflow218- User says "ship it", "deploy this", "publish this", "launch this", "get this out the door" -> run the ship workflow219- User invokes `/autoresearch:debug` -> run the debug loop220- User says "find all bugs", "hunt bugs", "debug this", "why is this failing", "investigate" -> run the debug loop221- User invokes `/autoresearch:fix` -> run the fix loop222- User says "fix all errors", "make tests pass", "fix the build", "clean up errors" -> run the fix loop223- User says "work autonomously", "iterate until done", "keep improving", "run overnight" -> run the loop224- Any task requiring repeated iteration cycles with measurable outcomes -> run the loop225226## Bounded Iterations227228By default, autoresearch loops **forever** until manually interrupted. To run exactly N iterations, add `Iterations: N` to your inline config.229230**Unlimited (default):**231```232/autoresearch233Goal: Increase test coverage to 90%234```235236**Bounded (N iterations):**237```238/autoresearch239Goal: Increase test coverage to 90%240Iterations: 25241```242243After N iterations Claude stops and prints a final summary with baseline -> current best, keeps/discards/crashes. If the goal is achieved before N iterations, Claude prints early completion and stops.244245### When to Use Bounded Iterations246247| Scenario | Recommendation |248|----------|---------------|249| Run overnight, review in morning | Unlimited (default) |250| Quick 30-min improvement session | `Iterations: 10` |251| Targeted fix with known scope | `Iterations: 5` |252| Exploratory -- see if approach works | `Iterations: 15` |253| CI/CD pipeline integration | `--iterations N` flag (set N based on time budget) |254255## Setup Phase (Do Once)256257**If the user provides Goal, Scope, Metric, and Verify inline** -> extract them and proceed to step 5.258259**If any critical field is missing** -> use `AskUserQuestion` to collect them interactively:260261### Interactive Setup (when invoked without full config)262263Scan the codebase first for smart defaults, then ask ALL questions in batched `AskUserQuestion` calls (max 4 per call). This gives users full clarity upfront.264265**Batch 1 -- Core config (4 questions in one call):**266267Use a SINGLE `AskUserQuestion` call with these 4 questions:268269| # | Header | Question | Options (smart defaults from codebase scan) |270|---|--------|----------|----------------------------------------------|271| 1 | `Goal` | "What do you want to improve?" | "Test coverage (higher)", "Bundle size (lower)", "Performance (faster)", "Code quality (fewer errors)" |272| 2 | `Scope` | "Which files can autoresearch modify?" | Suggested globs from project structure (e.g. "src/**/*.ts", "content/**/*.md") |273| 3 | `Metric` | "What number tells you if it got better? (must be a command output, not subjective)" | Detected options: "coverage % (higher)", "bundle size KB (lower)", "error count (lower)", "test pass count (higher)" |274| 4 | `Direction` | "Higher or lower is better?" | "Higher is better", "Lower is better" |275276**Batch 2 -- Verify + Guard + Launch (3 questions in one call):**277278| # | Header | Question | Options |279|---|--------|----------|---------|280| 5 | `Verify` | "What command produces the metric? (I'll dry-run it to confirm)" | Suggested commands from detected tooling |281| 6 | `Guard` | "Any command that must ALWAYS pass? (prevents regressions)" | "npm test", "tsc --noEmit", "npm run build", "Skip -- no guard" |282| 7 | `Launch` | "Ready to go?" | "Launch (unlimited)", "Launch with iteration limit", "Edit config", "Cancel" |283284**After Batch 2:** Dry-run the verify command. If it fails, ask user to fix or choose a different command. If it passes, proceed with launch choice.285286**IMPORTANT:** Always batch questions -- never ask one at a time. Users should see all config choices together for full context.287288### Setup Steps (after config is complete)2892901. **Read all in-scope files** for full context before any modification2912. **Define the goal** -- extracted from user input or inline config2923. **Define scope constraints** -- validated file globs2934. **Define guard (optional)** -- regression prevention command2945. **Create a results log** -- Track every iteration (see `references/results-logging.md`)2956. **Establish baseline** -- Run verification on current state AND guard (if set). Record as iteration #02967. **Confirm and go** -- Show user the setup, get confirmation, then BEGIN THE LOOP297298## The Loop299300Read `references/autonomous-loop-protocol.md` for full protocol details.301302```303LOOP (FOREVER or N times):304 1. Review: Read current state + git history + results log305 2. Ideate: Pick next change based on goal, past results, what hasn't been tried306 3. Modify: Make ONE focused change to in-scope files307 4. Commit: Git commit the change (before verification)308 5. Verify: Run the mechanical metric (tests, build, benchmark, etc.)309 6. Guard: If guard is set, run the guard command310 7. Decide:311 - IMPROVED + guard passed (or no guard) -> Keep commit, log "keep", advance312 - IMPROVED + guard FAILED -> Revert, then try to rework the optimization313 (max 2 attempts) so it improves the metric WITHOUT breaking the guard.314 Never modify guard/test files -- adapt the implementation instead.315 If still failing -> log "discard (guard failed)" and move on316 - SAME/WORSE -> Git revert, log "discard"317 - CRASHED -> Try to fix (max 3 attempts), else log "crash" and move on318 8. Log: Record result in results log319 9. Repeat: Go to step 1.320 - If unbounded: NEVER STOP. NEVER ASK "should I continue?"321 - If bounded (N): Stop after N iterations, print final summary322```323324## Critical Rules3253261. **Loop until done** -- Unbounded: loop until interrupted. Bounded: loop N times then summarize.3272. **Read before write** -- Always understand full context before modifying3283. **One change per iteration** -- Atomic changes. If it breaks, you know exactly why3294. **Mechanical verification only** -- No subjective "looks good". Use metrics3305. **Automatic rollback** -- Failed changes revert instantly. No debates3316. **Simplicity wins** -- Equal results + less code = KEEP. Tiny improvement + ugly complexity = DISCARD3327. **Git is memory** -- Every kept change committed. Agent reads history to learn patterns3338. **When stuck, think harder** -- Re-read files, re-read goal, combine near-misses, try radical changes. Don't ask for help unless truly blocked by missing access/permissions334335## Principles Reference336337See `references/core-principles.md` for the 7 generalizable principles from autoresearch.338339## Adapting to Different Domains340341| Domain | Metric | Scope | Verify Command | Guard |342|--------|--------|-------|----------------|-------|343| Backend code | Tests pass + coverage % | `src/**/*.ts` | `npm test` | -- |344| Frontend UI | Lighthouse score | `src/components/**` | `npx lighthouse` | `npm test` |345| ML training | val_bpb / loss | `train.py` | `uv run train.py` | -- |346| Blog/content | Word count + readability | `content/*.md` | Custom script | -- |347| Performance | Benchmark time (ms) | Target files | `npm run bench` | `npm test` |348| Refactoring | Tests pass + LOC reduced | Target module | `npm test && wc -l` | `npm run typecheck` |349| Security | OWASP + STRIDE coverage + findings | API/auth/middleware | `/autoresearch:security` | -- |350| Shipping | Checklist pass rate (%) | Any artifact | `/autoresearch:ship` | Domain-specific |351| Debugging | Bugs found + coverage | Target files | `/autoresearch:debug` | -- |352| Fixing | Error count (lower) | Target files | `/autoresearch:fix` | `npm test` |353354Adapt the loop to your domain. The PRINCIPLES are universal; the METRICS are domain-specific.