Customization
Before executing, check for user customizations at:
~/.claude/PAI/USER/SKILLCUSTOMIZATIONS/Evals/
If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.
🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)
You MUST send this notification BEFORE doing anything else when this skill is invoked.
Send voice notification:
curl -s -X POST http://localhost:8888/notify \
-H "Content-Type: application/json" \
-d '{"message": "Running the WORKFLOWNAME workflow in the Evals skill to ACTION"}' \
> /dev/null 2>&1 &
Output text notification:
Running the **WorkflowName** workflow in the **Evals** skill to ACTION...
This is not optional. Execute this curl command immediately upon skill invocation.
Evals - AI Agent Evaluation Framework
Comprehensive agent evaluation system based on Anthropic's "Demystifying Evals for AI Agents" (Jan 2026).
Key differentiator: Evaluates agent workflows (transcripts, tool calls, multi-turn conversations), not just single outputs.
When to Activate
- "run evals", "test this agent", "evaluate", "check quality", "benchmark"
- "regression test", "capability test"
- Compare agent behaviors across changes
- Validate agent workflows before deployment
- Verify ALGORITHM ISC rows
- Create new evaluation tasks from failures
Core Concepts
Three Grader Types
| Type |
Strengths |
Weaknesses |
Use For |
| Code-based |
Fast, cheap, deterministic, reproducible |
Brittle, lacks nuance |
Tests, state checks, tool verification |
| Model-based |
Flexible, captures nuance, scalable |
Non-deterministic, expensive |
Quality rubrics, assertions, comparisons |
| Human |
Gold standard, handles subjectivity |
Expensive, slow |
Calibration, spot checks, A/B testing |
Evaluation Types
| Type |
Pass Target |
Purpose |
| Capability |
~70% |
Stretch goals, measuring improvement potential |
| Regression |
~99% |
Quality gates, detecting backsliding |
Key Metrics
- pass@k: Probability of at least 1 success in k trials (measures capability)
- pass^k: Probability all k trials succeed (measures consistency/reliability)
Workflow Routing
| Request Pattern |
Route To |
| Run eval, evaluate suite, run tests, benchmark |
Workflows/RunEval.md |
| Compare models, model comparison, A/B test models |
Workflows/CompareModels.md |
| Compare prompts, prompt comparison, test prompts |
Workflows/ComparePrompts.md |
| Create judge, model grader, evaluation judge |
Workflows/CreateJudge.md |
| Create use case, new eval, test case, create suite |
Workflows/CreateUseCase.md |
| View results, eval results, scores, pass rate |
Workflows/ViewResults.md |
CLI Quick Reference
| Trigger |
Tool |
| Run suite |
Tools/AlgorithmBridge.ts |
| Log failure |
Tools/FailureToTask.ts log |
| Convert failures |
Tools/FailureToTask.ts convert-all |
| Create suite |
Tools/SuiteManager.ts create |
| Check saturation |
Tools/SuiteManager.ts check-saturation |
Quick Reference
CLI Commands
# Run an eval suite
bun run ~/.claude/skills/Utilities/Evals/Tools/AlgorithmBridge.ts -s <suite>
# Log a failure for later conversion
bun run ~/.claude/skills/Utilities/Evals/Tools/FailureToTask.ts log "description" -c category -s severity
# Convert failures to test tasks
bun run ~/.claude/skills/Utilities/Evals/Tools/FailureToTask.ts convert-all
# Manage suites
bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts create <name> -t capability -d "description"
bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts list
bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts check-saturation <name>
bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts graduate <name>
ALGORITHM Integration
Evals is a verification method for THE ALGORITHM ISC rows:
# Run eval and update ISC row
bun run ~/.claude/skills/Utilities/Evals/Tools/AlgorithmBridge.ts -s regression-core -r 3 -u
ISC rows can specify eval verification:
| # | What Ideal Looks Like | Verify |
|---|----------------------|--------|
| 1 | Auth bypass fixed | eval:auth-security |
| 2 | Tests all pass | eval:regression |
Available Graders
Code-Based (Fast, Deterministic)
| Grader |
Use Case |
string_match |
Exact substring matching |
regex_match |
Pattern matching |
binary_tests |
Run test files |
static_analysis |
Lint, type-check, security scan |
state_check |
Verify system state after execution |
tool_calls |
Verify specific tools were called |
Model-Based (Nuanced)
| Grader |
Use Case |
llm_rubric |
Score against detailed rubric |
natural_language_assert |
Check assertions are true |
pairwise_comparison |
Compare to reference with position swap |
Domain Patterns
Pre-configured grader stacks for common agent types:
| Domain |
Primary Graders |
coding |
binary_tests + static_analysis + tool_calls + llm_rubric |
conversational |
llm_rubric + natural_language_assert + state_check |
research |
llm_rubric + natural_language_assert + tool_calls |
computer_use |
state_check + tool_calls + llm_rubric |
See Data/DomainPatterns.yaml for full configurations.
Task Schema (YAML)
task:
id: "fix-auth-bypass_1"
description: "Fix authentication bypass when password is empty"
type: regression # or capability
domain: coding
graders:
- type: binary_tests
required: [test_empty_pw.py]
weight: 0.30
- type: tool_calls
weight: 0.20
params:
sequence: [read_file, edit_file, run_tests]
- type: llm_rubric
weight: 0.50
params:
rubric: prompts/security_review.md
trials: 3
pass_threshold: 0.75
Resource Index
| Resource |
Purpose |
Types/index.ts |
Core type definitions |
Graders/CodeBased/ |
Deterministic graders |
Graders/ModelBased/ |
LLM-powered graders |
Tools/TranscriptCapture.ts |
Capture agent trajectories |
Tools/TrialRunner.ts |
Multi-trial execution with pass@k |
Tools/SuiteManager.ts |
Suite management and saturation |
Tools/FailureToTask.ts |
Convert failures to test tasks |
Tools/AlgorithmBridge.ts |
ALGORITHM integration |
Data/DomainPatterns.yaml |
Domain-specific grader configs |
Key Principles (from Anthropic)
- Start with 20-50 real failures - Don't overthink, capture what actually broke
- Unambiguous tasks - Two experts should reach identical verdicts
- Balanced problem sets - Test both "should do" AND "should NOT do"
- Grade outputs, not paths - Don't penalize valid creative solutions
- Calibrate LLM judges - Against human expert judgment
- Check transcripts regularly - Verify graders work correctly
- Monitor saturation - Graduate to regression when hitting 95%+
- Build infrastructure early - Evals shape how quickly you can adopt new models
Related
- ALGORITHM: Evals is a verification method
- Science: Evals implements scientific method
- Browser: For visual verification graders
1---2name: evals3description: Objective eval metrics via code/model/human graders with pass@k/pass^k scoring. USE WHEN eval, evaluate, test agent, benchmark, verify behavior, regression test, capability test, run eval, compare models, compare prompts, create judge, create use case, view results, failure to task, suite manager, transcript capture, trial runner.4---56## Customization78**Before executing, check for user customizations at:**9`~/.claude/PAI/USER/SKILLCUSTOMIZATIONS/Evals/`1011If this directory exists, load and apply any PREFERENCES.md, configurations, or resources found there. These override default behavior. If the directory does not exist, proceed with skill defaults.121314## 🚨 MANDATORY: Voice Notification (REQUIRED BEFORE ANY ACTION)1516**You MUST send this notification BEFORE doing anything else when this skill is invoked.**17181. **Send voice notification**:19 ```bash20 curl -s -X POST http://localhost:8888/notify \21 -H "Content-Type: application/json" \22 -d '{"message": "Running the WORKFLOWNAME workflow in the Evals skill to ACTION"}' \23 > /dev/null 2>&1 &24 ```25262. **Output text notification**:27 ```28 Running the **WorkflowName** workflow in the **Evals** skill to ACTION...29 ```3031**This is not optional. Execute this curl command immediately upon skill invocation.**3233# Evals - AI Agent Evaluation Framework3435Comprehensive agent evaluation system based on Anthropic's "Demystifying Evals for AI Agents" (Jan 2026).3637**Key differentiator:** Evaluates agent *workflows* (transcripts, tool calls, multi-turn conversations), not just single outputs.3839---4041## When to Activate4243- "run evals", "test this agent", "evaluate", "check quality", "benchmark"44- "regression test", "capability test"45- Compare agent behaviors across changes46- Validate agent workflows before deployment47- Verify ALGORITHM ISC rows48- Create new evaluation tasks from failures4950---5152## Core Concepts5354### Three Grader Types5556| Type | Strengths | Weaknesses | Use For |57|------|-----------|------------|---------|58| **Code-based** | Fast, cheap, deterministic, reproducible | Brittle, lacks nuance | Tests, state checks, tool verification |59| **Model-based** | Flexible, captures nuance, scalable | Non-deterministic, expensive | Quality rubrics, assertions, comparisons |60| **Human** | Gold standard, handles subjectivity | Expensive, slow | Calibration, spot checks, A/B testing |6162### Evaluation Types6364| Type | Pass Target | Purpose |65|------|-------------|---------|66| **Capability** | ~70% | Stretch goals, measuring improvement potential |67| **Regression** | ~99% | Quality gates, detecting backsliding |6869### Key Metrics7071- **pass@k**: Probability of at least 1 success in k trials (measures capability)72- **pass^k**: Probability all k trials succeed (measures consistency/reliability)7374---7576## Workflow Routing7778| Request Pattern | Route To |79|---|---|80| Run eval, evaluate suite, run tests, benchmark | `Workflows/RunEval.md` |81| Compare models, model comparison, A/B test models | `Workflows/CompareModels.md` |82| Compare prompts, prompt comparison, test prompts | `Workflows/ComparePrompts.md` |83| Create judge, model grader, evaluation judge | `Workflows/CreateJudge.md` |84| Create use case, new eval, test case, create suite | `Workflows/CreateUseCase.md` |85| View results, eval results, scores, pass rate | `Workflows/ViewResults.md` |8687### CLI Quick Reference8889| Trigger | Tool |90|---------|------|91| Run suite | `Tools/AlgorithmBridge.ts` |92| Log failure | `Tools/FailureToTask.ts log` |93| Convert failures | `Tools/FailureToTask.ts convert-all` |94| Create suite | `Tools/SuiteManager.ts create` |95| Check saturation | `Tools/SuiteManager.ts check-saturation` |9697---9899## Quick Reference100101### CLI Commands102103```bash104# Run an eval suite105bun run ~/.claude/skills/Utilities/Evals/Tools/AlgorithmBridge.ts -s <suite>106107# Log a failure for later conversion108bun run ~/.claude/skills/Utilities/Evals/Tools/FailureToTask.ts log "description" -c category -s severity109110# Convert failures to test tasks111bun run ~/.claude/skills/Utilities/Evals/Tools/FailureToTask.ts convert-all112113# Manage suites114bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts create <name> -t capability -d "description"115bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts list116bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts check-saturation <name>117bun run ~/.claude/skills/Utilities/Evals/Tools/SuiteManager.ts graduate <name>118```119120### ALGORITHM Integration121122Evals is a verification method for THE ALGORITHM ISC rows:123124```bash125# Run eval and update ISC row126bun run ~/.claude/skills/Utilities/Evals/Tools/AlgorithmBridge.ts -s regression-core -r 3 -u127```128129ISC rows can specify eval verification:130```131| # | What Ideal Looks Like | Verify |132|---|----------------------|--------|133| 1 | Auth bypass fixed | eval:auth-security |134| 2 | Tests all pass | eval:regression |135```136137---138139## Available Graders140141### Code-Based (Fast, Deterministic)142143| Grader | Use Case |144|--------|----------|145| `string_match` | Exact substring matching |146| `regex_match` | Pattern matching |147| `binary_tests` | Run test files |148| `static_analysis` | Lint, type-check, security scan |149| `state_check` | Verify system state after execution |150| `tool_calls` | Verify specific tools were called |151152### Model-Based (Nuanced)153154| Grader | Use Case |155|--------|----------|156| `llm_rubric` | Score against detailed rubric |157| `natural_language_assert` | Check assertions are true |158| `pairwise_comparison` | Compare to reference with position swap |159160---161162## Domain Patterns163164Pre-configured grader stacks for common agent types:165166| Domain | Primary Graders |167|--------|-----------------|168| `coding` | binary_tests + static_analysis + tool_calls + llm_rubric |169| `conversational` | llm_rubric + natural_language_assert + state_check |170| `research` | llm_rubric + natural_language_assert + tool_calls |171| `computer_use` | state_check + tool_calls + llm_rubric |172173See `Data/DomainPatterns.yaml` for full configurations.174175---176177## Task Schema (YAML)178179```yaml180task:181 id: "fix-auth-bypass_1"182 description: "Fix authentication bypass when password is empty"183 type: regression # or capability184 domain: coding185186 graders:187 - type: binary_tests188 required: [test_empty_pw.py]189 weight: 0.30190191 - type: tool_calls192 weight: 0.20193 params:194 sequence: [read_file, edit_file, run_tests]195196 - type: llm_rubric197 weight: 0.50198 params:199 rubric: prompts/security_review.md200201 trials: 3202 pass_threshold: 0.75203```204205---206207## Resource Index208209| Resource | Purpose |210|----------|---------|211| `Types/index.ts` | Core type definitions |212| `Graders/CodeBased/` | Deterministic graders |213| `Graders/ModelBased/` | LLM-powered graders |214| `Tools/TranscriptCapture.ts` | Capture agent trajectories |215| `Tools/TrialRunner.ts` | Multi-trial execution with pass@k |216| `Tools/SuiteManager.ts` | Suite management and saturation |217| `Tools/FailureToTask.ts` | Convert failures to test tasks |218| `Tools/AlgorithmBridge.ts` | ALGORITHM integration |219| `Data/DomainPatterns.yaml` | Domain-specific grader configs |220221---222223## Key Principles (from Anthropic)2242251. **Start with 20-50 real failures** - Don't overthink, capture what actually broke2262. **Unambiguous tasks** - Two experts should reach identical verdicts2273. **Balanced problem sets** - Test both "should do" AND "should NOT do"2284. **Grade outputs, not paths** - Don't penalize valid creative solutions2295. **Calibrate LLM judges** - Against human expert judgment2306. **Check transcripts regularly** - Verify graders work correctly2317. **Monitor saturation** - Graduate to regression when hitting 95%+2328. **Build infrastructure early** - Evals shape how quickly you can adopt new models233234---235236## Related237238- **ALGORITHM**: Evals is a verification method239- **Science**: Evals implements scientific method240- **Browser**: For visual verification graders