Prompt Optimizer
Optimize prompts with evals. Keep every instruction, example, and external context reference causal.
Load Only What You Need
| Need |
Read |
| New prompt |
references/core-patterns.md, references/model-family-notes.md, references/transformed-examples.md |
| Existing prompt |
references/meta-optimization-loop.md, references/core-patterns.md, references/model-family-notes.md |
| Model-family port |
references/model-family-notes.md, references/core-patterns.md |
| Repeated failures |
references/meta-optimization-loop.md, references/core-patterns.md |
| Weak or ambiguous draft |
references/transformed-examples.md |
| Provenance |
SOURCES.md |
Step 1: Capture Contract
Record before editing:
- task type: new, refine, port, or debug
- target model family and snapshot, if known
- prompt surface:
system, developer, user, tool descriptions, examples, schemas
- layer owners: platform, deployer/persona, retrieved context, user payload
- objective and non-goals
- inputs, tools, and external files available
- required output shape
- success criteria and failure cases
- hard constraints: latency, verbosity, safety, budget, tool use, style
If success criteria or examples are missing, create a small eval set first.
If the bottleneck is model choice, retrieval, tool schema, or missing evals, say so before rewriting.
Step 2: Inventory External Context
For repo or agent prompts, list stable context by exact path:
| Context type |
Examples |
| Agent rules |
AGENTS.md, CLAUDE.md |
| Specs |
specs/*.md, docs/api.md |
| Policies |
SECURITY.md, docs/releasing.md |
| Examples |
examples/, tests/fixtures/ |
Rules:
- Reference stable files by repo-relative path instead of copying them.
- Paste only excerpts needed for the prompt or eval case.
- Mark whether a file is
loaded, referenced, or out of scope.
- Avoid vague context pointers such as "read the docs".
Step 3: Choose Model Strategy
Read references/model-family-notes.md.
- Known family: optimize for that family.
- Unknown family: write a portable base plus short adapter notes.
- Snapshot changes: rerun evals.
- Cross-family divergence: specialize only the failing layer.
Step 4: Shape Prompt
Read references/core-patterns.md.
- Put stable policy in
system or developer.
- Put task-local facts, retrieved context, and variables in user-facing sections.
- Keep one owner per behavior rule.
- Use headings or tags only to separate content types.
- Put tool policy in prompt text; keep schemas in provider-native tools.
- Keep persona light unless it changes behavior.
- Use the shortest wording that preserves the constraint.
- Cut filler, repeated reminders, dead examples, and rationale that does not affect evals.
Step 5: Optimize
Read references/meta-optimization-loop.md for refinements.
- Baseline the current prompt on the same eval slice.
- Cluster failures by root cause.
- Write concrete edit criticisms.
- Generate two to four candidates:
- minimal-diff repair
- structure-first rewrite
- examples-first or tool-rule variant
- provider adapter when needed
- Compare candidates on the same cases.
- Keep a short optimization log.
- Validate the winner on holdout cases.
- Stop on plateau, oscillation, overfit, excessive cost, or non-prompt bottleneck.
Step 6: Return Package
Return:
Target
Success Criteria
External Context
Optimized Prompt
Adapter Notes
Eval Set
Optimization Log
Residual Risks
For existing prompts, include a concise diff-style note of the main behavioral changes.
Failure Modes
- editing before defining the eval target
- mixing policy, examples, and raw context without boundaries
- duplicating rules across layers
- putting durable policy in user payloads
- asking for chain-of-thought
- keeping contradictory legacy instructions
- overfitting to one or two examples
- retaining examples that no longer improve evals
- fixing tool-use failures only in prompt text when tool descriptions or schemas are weak
- adding markup that does not reduce ambiguity
- using persona as a substitute for behavior rules
1---2name: prompt-optimizer3description: Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates with evals.4---56# Prompt Optimizer78Optimize prompts with evals. Keep every instruction, example, and external context reference causal.910## Load Only What You Need1112| Need | Read |13|------|------|14| New prompt | `references/core-patterns.md`, `references/model-family-notes.md`, `references/transformed-examples.md` |15| Existing prompt | `references/meta-optimization-loop.md`, `references/core-patterns.md`, `references/model-family-notes.md` |16| Model-family port | `references/model-family-notes.md`, `references/core-patterns.md` |17| Repeated failures | `references/meta-optimization-loop.md`, `references/core-patterns.md` |18| Weak or ambiguous draft | `references/transformed-examples.md` |19| Provenance | `SOURCES.md` |2021## Step 1: Capture Contract2223Record before editing:2425- task type: new, refine, port, or debug26- target model family and snapshot, if known27- prompt surface: `system`, `developer`, `user`, tool descriptions, examples, schemas28- layer owners: platform, deployer/persona, retrieved context, user payload29- objective and non-goals30- inputs, tools, and external files available31- required output shape32- success criteria and failure cases33- hard constraints: latency, verbosity, safety, budget, tool use, style3435If success criteria or examples are missing, create a small eval set first.36If the bottleneck is model choice, retrieval, tool schema, or missing evals, say so before rewriting.3738## Step 2: Inventory External Context3940For repo or agent prompts, list stable context by exact path:4142| Context type | Examples |43|--------------|----------|44| Agent rules | `AGENTS.md`, `CLAUDE.md` |45| Specs | `specs/*.md`, `docs/api.md` |46| Policies | `SECURITY.md`, `docs/releasing.md` |47| Examples | `examples/`, `tests/fixtures/` |4849Rules:5051- Reference stable files by repo-relative path instead of copying them.52- Paste only excerpts needed for the prompt or eval case.53- Mark whether a file is `loaded`, `referenced`, or `out of scope`.54- Avoid vague context pointers such as "read the docs".5556## Step 3: Choose Model Strategy5758Read `references/model-family-notes.md`.5960- Known family: optimize for that family.61- Unknown family: write a portable base plus short adapter notes.62- Snapshot changes: rerun evals.63- Cross-family divergence: specialize only the failing layer.6465## Step 4: Shape Prompt6667Read `references/core-patterns.md`.6869- Put stable policy in `system` or `developer`.70- Put task-local facts, retrieved context, and variables in user-facing sections.71- Keep one owner per behavior rule.72- Use headings or tags only to separate content types.73- Put tool policy in prompt text; keep schemas in provider-native tools.74- Keep persona light unless it changes behavior.75- Use the shortest wording that preserves the constraint.76- Cut filler, repeated reminders, dead examples, and rationale that does not affect evals.7778## Step 5: Optimize7980Read `references/meta-optimization-loop.md` for refinements.81821. Baseline the current prompt on the same eval slice.832. Cluster failures by root cause.843. Write concrete edit criticisms.854. Generate two to four candidates:86 - minimal-diff repair87 - structure-first rewrite88 - examples-first or tool-rule variant89 - provider adapter when needed905. Compare candidates on the same cases.916. Keep a short optimization log.927. Validate the winner on holdout cases.938. Stop on plateau, oscillation, overfit, excessive cost, or non-prompt bottleneck.9495## Step 6: Return Package9697Return:98991. `Target`1002. `Success Criteria`1013. `External Context`1024. `Optimized Prompt`1035. `Adapter Notes`1046. `Eval Set`1057. `Optimization Log`1068. `Residual Risks`107108For existing prompts, include a concise diff-style note of the main behavioral changes.109110## Failure Modes111112- editing before defining the eval target113- mixing policy, examples, and raw context without boundaries114- duplicating rules across layers115- putting durable policy in user payloads116- asking for chain-of-thought117- keeping contradictory legacy instructions118- overfitting to one or two examples119- retaining examples that no longer improve evals120- fixing tool-use failures only in prompt text when tool descriptions or schemas are weak121- adding markup that does not reduce ambiguity122- using persona as a substitute for behavior rules