Evals - AI Agent Evaluation Framework
Comprehensive agent evaluation system based on Anthropic's "Demystifying Evals for AI Agents" (Jan 2026).
Key differentiator: Evaluates agent workflows (transcripts, tool calls, multi-turn conversations), not just single outputs.
When to Activate
- "run evals", "test this agent", "evaluate", "check quality", "benchmark"
- "regression test", "capability test"
- Compare agent behaviors across changes
- Validate agent workflows before deployment
- Verify ALGORITHM ISC rows
- Create new evaluation tasks from failures
Core Concepts
Three Grader Types
| Type |
Strengths |
Weaknesses |
Use For |
| Code-based |
Fast, cheap, deterministic, reproducible |
Brittle, lacks nuance |
Tests, state checks, tool verification |
| Model-based |
Flexible, captures nuance, scalable |
Non-deterministic, expensive |
Quality rubrics, assertions, comparisons |
| Human |
Gold standard, handles subjectivity |
Expensive, slow |
Calibration, spot checks, A/B testing |
Evaluation Types
| Type |
Pass Target |
Purpose |
| Capability |
~70% |
Stretch goals, measuring improvement potential |
| Regression |
~99% |
Quality gates, detecting backsliding |
Key Metrics
- pass@k: Probability of at least 1 success in k trials (measures capability)
- pass^k: Probability all k trials succeed (measures consistency/reliability)
Workflow Routing
| Trigger |
Workflow |
| "run evals", "evaluate suite" |
Run suite via Tools/AlgorithmBridge.ts |
| "log failure" |
Log failure via Tools/FailureToTask.ts log |
| "convert failures" |
Convert to tasks via Tools/FailureToTask.ts convert-all |
| "create suite" |
Create suite via Tools/SuiteManager.ts create |
| "check saturation" |
Check via Tools/SuiteManager.ts check-saturation |
Quick Reference
CLI Commands
# Run an eval suite
bun run ./Tools/AlgorithmBridge.ts -s <suite>
# Log a failure for later conversion
bun run ./Tools/FailureToTask.ts log "description" -c category -s severity
# Convert failures to test tasks
bun run ./Tools/FailureToTask.ts convert-all
# Manage suites
bun run ./Tools/SuiteManager.ts create <name> -t capability -d "description"
bun run ./Tools/SuiteManager.ts list
bun run ./Tools/SuiteManager.ts check-saturation <name>
bun run ./Tools/SuiteManager.ts graduate <name>
ALGORITHM Integration
Evals is a verification method for THE ALGORITHM ISC rows:
# Run eval and update ISC row
bun run ./Tools/AlgorithmBridge.ts -s regression-core -r 3 -u
ISC rows can specify eval verification:
| # | What Ideal Looks Like | Verify |
|---|----------------------|--------|
| 1 | Auth bypass fixed | eval:auth-security |
| 2 | Tests all pass | eval:regression |
Available Graders
Code-Based (Fast, Deterministic)
| Grader |
Use Case |
string_match |
Exact substring matching |
regex_match |
Pattern matching |
binary_tests |
Run test files |
static_analysis |
Lint, type-check, security scan |
state_check |
Verify system state after execution |
tool_calls |
Verify specific tools were called |
Model-Based (Nuanced)
| Grader |
Use Case |
llm_rubric |
Score against detailed rubric |
natural_language_assert |
Check assertions are true |
pairwise_comparison |
Compare to reference with position swap |
Domain Patterns
Pre-configured grader stacks for common agent types:
| Domain |
Primary Graders |
coding |
binary_tests + static_analysis + tool_calls + llm_rubric |
conversational |
llm_rubric + natural_language_assert + state_check |
research |
llm_rubric + natural_language_assert + tool_calls |
computer_use |
state_check + tool_calls + llm_rubric |
See Data/DomainPatterns.yaml for full configurations.
Task Schema (YAML)
task:
id: "fix-auth-bypass_1"
description: "Fix authentication bypass when password is empty"
type: regression # or capability
domain: coding
graders:
- type: binary_tests
required: [test_empty_pw.py]
weight: 0.30
- type: tool_calls
weight: 0.20
params:
sequence: [read_file, edit_file, run_tests]
- type: llm_rubric
weight: 0.50
params:
rubric: prompts/security_review.md
trials: 3
pass_threshold: 0.75
Resource Index
| Resource |
Purpose |
Types/index.ts |
Core type definitions |
Graders/CodeBased/ |
Deterministic graders |
Graders/ModelBased/ |
LLM-powered graders |
Tools/TranscriptCapture.ts |
Capture agent trajectories |
Tools/TrialRunner.ts |
Multi-trial execution with pass@k |
Tools/SuiteManager.ts |
Suite management and saturation |
Tools/FailureToTask.ts |
Convert failures to test tasks |
Tools/AlgorithmBridge.ts |
ALGORITHM integration |
Data/DomainPatterns.yaml |
Domain-specific grader configs |
Key Principles (from Anthropic)
- Start with 20-50 real failures - Don't overthink, capture what actually broke
- Unambiguous tasks - Two experts should reach identical verdicts
- Balanced problem sets - Test both "should do" AND "should NOT do"
- Grade outputs, not paths - Don't penalize valid creative solutions
- Calibrate LLM judges - Against human expert judgment
- Check transcripts regularly - Verify graders work correctly
- Monitor saturation - Graduate to regression when hitting 95%+
- Build infrastructure early - Evals shape how quickly you can adopt new models
Related
- ALGORITHM: Evals is a verification method
- Science: Evals implements scientific method
- Browser: For visual verification graders
1---2name: evals3description: Agent evaluation framework based on Anthropic's best practices. USE WHEN eval, evaluate, test agent, benchmark, verify behavior, regression test, capability test. Includes three grader types (code-based, model-based, human), transcript capture, pass@k/pass^k metrics, and ALGORITHM integration.4---56# Evals - AI Agent Evaluation Framework78Comprehensive agent evaluation system based on Anthropic's "Demystifying Evals for AI Agents" (Jan 2026).910**Key differentiator:** Evaluates agent *workflows* (transcripts, tool calls, multi-turn conversations), not just single outputs.1112---1314## When to Activate1516- "run evals", "test this agent", "evaluate", "check quality", "benchmark"17- "regression test", "capability test"18- Compare agent behaviors across changes19- Validate agent workflows before deployment20- Verify ALGORITHM ISC rows21- Create new evaluation tasks from failures2223---2425## Core Concepts2627### Three Grader Types2829| Type | Strengths | Weaknesses | Use For |30|------|-----------|------------|---------|31| **Code-based** | Fast, cheap, deterministic, reproducible | Brittle, lacks nuance | Tests, state checks, tool verification |32| **Model-based** | Flexible, captures nuance, scalable | Non-deterministic, expensive | Quality rubrics, assertions, comparisons |33| **Human** | Gold standard, handles subjectivity | Expensive, slow | Calibration, spot checks, A/B testing |3435### Evaluation Types3637| Type | Pass Target | Purpose |38|------|-------------|---------|39| **Capability** | ~70% | Stretch goals, measuring improvement potential |40| **Regression** | ~99% | Quality gates, detecting backsliding |4142### Key Metrics4344- **pass@k**: Probability of at least 1 success in k trials (measures capability)45- **pass^k**: Probability all k trials succeed (measures consistency/reliability)4647---4849## Workflow Routing5051| Trigger | Workflow |52|---------|----------|53| "run evals", "evaluate suite" | Run suite via `Tools/AlgorithmBridge.ts` |54| "log failure" | Log failure via `Tools/FailureToTask.ts log` |55| "convert failures" | Convert to tasks via `Tools/FailureToTask.ts convert-all` |56| "create suite" | Create suite via `Tools/SuiteManager.ts create` |57| "check saturation" | Check via `Tools/SuiteManager.ts check-saturation` |5859---6061## Quick Reference6263### CLI Commands6465```bash66# Run an eval suite67bun run ./Tools/AlgorithmBridge.ts -s <suite>6869# Log a failure for later conversion70bun run ./Tools/FailureToTask.ts log "description" -c category -s severity7172# Convert failures to test tasks73bun run ./Tools/FailureToTask.ts convert-all7475# Manage suites76bun run ./Tools/SuiteManager.ts create <name> -t capability -d "description"77bun run ./Tools/SuiteManager.ts list78bun run ./Tools/SuiteManager.ts check-saturation <name>79bun run ./Tools/SuiteManager.ts graduate <name>80```8182### ALGORITHM Integration8384Evals is a verification method for THE ALGORITHM ISC rows:8586```bash87# Run eval and update ISC row88bun run ./Tools/AlgorithmBridge.ts -s regression-core -r 3 -u89```9091ISC rows can specify eval verification:92```93| # | What Ideal Looks Like | Verify |94|---|----------------------|--------|95| 1 | Auth bypass fixed | eval:auth-security |96| 2 | Tests all pass | eval:regression |97```9899---100101## Available Graders102103### Code-Based (Fast, Deterministic)104105| Grader | Use Case |106|--------|----------|107| `string_match` | Exact substring matching |108| `regex_match` | Pattern matching |109| `binary_tests` | Run test files |110| `static_analysis` | Lint, type-check, security scan |111| `state_check` | Verify system state after execution |112| `tool_calls` | Verify specific tools were called |113114### Model-Based (Nuanced)115116| Grader | Use Case |117|--------|----------|118| `llm_rubric` | Score against detailed rubric |119| `natural_language_assert` | Check assertions are true |120| `pairwise_comparison` | Compare to reference with position swap |121122---123124## Domain Patterns125126Pre-configured grader stacks for common agent types:127128| Domain | Primary Graders |129|--------|-----------------|130| `coding` | binary_tests + static_analysis + tool_calls + llm_rubric |131| `conversational` | llm_rubric + natural_language_assert + state_check |132| `research` | llm_rubric + natural_language_assert + tool_calls |133| `computer_use` | state_check + tool_calls + llm_rubric |134135See `Data/DomainPatterns.yaml` for full configurations.136137---138139## Task Schema (YAML)140141```yaml142task:143 id: "fix-auth-bypass_1"144 description: "Fix authentication bypass when password is empty"145 type: regression # or capability146 domain: coding147148 graders:149 - type: binary_tests150 required: [test_empty_pw.py]151 weight: 0.30152153 - type: tool_calls154 weight: 0.20155 params:156 sequence: [read_file, edit_file, run_tests]157158 - type: llm_rubric159 weight: 0.50160 params:161 rubric: prompts/security_review.md162163 trials: 3164 pass_threshold: 0.75165```166167---168169## Resource Index170171| Resource | Purpose |172|----------|---------|173| `Types/index.ts` | Core type definitions |174| `Graders/CodeBased/` | Deterministic graders |175| `Graders/ModelBased/` | LLM-powered graders |176| `Tools/TranscriptCapture.ts` | Capture agent trajectories |177| `Tools/TrialRunner.ts` | Multi-trial execution with pass@k |178| `Tools/SuiteManager.ts` | Suite management and saturation |179| `Tools/FailureToTask.ts` | Convert failures to test tasks |180| `Tools/AlgorithmBridge.ts` | ALGORITHM integration |181| `Data/DomainPatterns.yaml` | Domain-specific grader configs |182183---184185## Key Principles (from Anthropic)1861871. **Start with 20-50 real failures** - Don't overthink, capture what actually broke1882. **Unambiguous tasks** - Two experts should reach identical verdicts1893. **Balanced problem sets** - Test both "should do" AND "should NOT do"1904. **Grade outputs, not paths** - Don't penalize valid creative solutions1915. **Calibrate LLM judges** - Against human expert judgment1926. **Check transcripts regularly** - Verify graders work correctly1937. **Monitor saturation** - Graduate to regression when hitting 95%+1948. **Build infrastructure early** - Evals shape how quickly you can adopt new models195196---197198## Related199200- **ALGORITHM**: Evals is a verification method201- **Science**: Evals implements scientific method202- **Browser**: For visual verification graders