Evals — Assertion-First AI Evaluation
What it is
An eval gives an AI an input, then applies assertions to its output to measure success (Anthropic's definition). A case is {id, prompt, assert:[...]}. Each assertion is either deterministic (code, fast/free) or model-graded (an LLM judge). Cases run multiple trials; we report pass^k (all trials pass — the honest metric for a reliability-critical agent) and pass@k (any trial passes). Everything routes through Inference.ts — subscription-billed, no API-key path, no external deps.
Grounded in Anthropic's current doctrine — Demystifying evals for AI agents, Define success criteria / develop tests, and the skill-creator {text, passed, evidence} assertion convention. The typed-assert layer is promptfoo-shaped but our own TS.
Freshness contract: "aligned to Anthropic's doctrine" is a live claim, not a snapshot. When designing a new suite class or touching the ## Doctrine section below, re-fetch the Demystifying-evals doc and flag where it has moved past what's encoded here. Advisory only — report divergence, never auto-adopt, and an unreachable URL never blocks a run.
The canonical path (v2)
| Tool |
Role |
Tools/Assertions.ts |
Deterministic assert engine: equals, contains, icontains, contains-all/any, regex, starts-with, ends-with, is-json, contains-json, max-length, min-length, each with not- negation. Sync, no model call. |
Tools/Judge.ts |
Model-graded asserts llm-rubric (1–5 → 0–1, threshold) and llm-assert (NL assertions → TRUE/FALSE/UNKNOWN). Forced-structured JSON verdict, reason-then-score, distinct judge level, Unknown→miss escape hatch. |
Tools/EvalRunner.ts |
Loads a suite, runs the agent-under-test per case (single-shot inference against the target system prompt), applies asserts, computes pass^k/pass@k, persists transcripts + latest.json. |
Tools/SuiteManager.ts |
Suite listing + saturation tracking. |
Tools/FailureToTask.ts |
Convert real failures into cases (seed from 20–50 real failures). |
# Run a suite (USER-customization suites resolve before the skill's own)
bun run ${LIFEOS_SKILL_DIR}/Tools/EvalRunner.ts -s <suite> [-t trials] [--json]
# Sanity-check the assert engine / judge
bun run ${LIFEOS_SKILL_DIR}/Tools/Assertions.ts # 16-case self-test
bun run ${LIFEOS_SKILL_DIR}/Tools/Judge.ts # good-vs-bad discrimination
Workflow Routing
| Workflow |
Trigger |
File |
| RunEval |
"run the eval", "run suite", "evaluate this", "grade output" |
Workflows/RunEval.md |
| CreateUseCase |
"new eval", "create a suite", "eval for X", "what should I test" |
Workflows/CreateUseCase.md |
| CreateJudge |
"write a judge", "llm-rubric", "grading criteria", "judge prompt" |
Workflows/CreateJudge.md |
| ComparePrompts |
"compare prompts", "which prompt is better", "A/B this prompt" |
Workflows/ComparePrompts.md |
| CompareModels |
"compare models", "which model is better", "is the cheaper rung enough" |
Workflows/CompareModels.md |
| ViewResults |
"eval results", "how did it score", "show the last run", "saturation" |
Workflows/ViewResults.md |
| CreateScenario |
"create a scenario", "multi-turn eval", "scenario test" |
Workflows/CreateScenario.md |
| RunScenario |
"run the scenario", "run multi-turn" |
Workflows/RunScenario.md |
Suite / case schema (assertion-first)
name: my-suite
type: regression # or capability
pass_threshold: 0.75
agent_level: medium # agent-under-test inference level
judge_level: high # judge != generator (Anthropic best practice)
trials: 3
# system_prompt: optional override; default = live system prompt + DA identity
cases:
- id: descriptive_name
prompt: "the user turn sent to the agent-under-test"
assert:
- type: not-contains # deterministic
value: "should work"
weight: 1
- type: llm-rubric # model-graded, weighted for partial credit
weight: 2
value: "Does the output tie any done-claim to verification evidence?"
- type: llm-assert
weight: 1
value: ["The output does not claim success without evidence"]
- id: should_not_case # balance: test should-do AND should-not
negative: true
prompt: "..."
assert: [...]
Identity-bound suites (e.g. {{DA_NAME}}'s dispositions) live in LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/Suites/ — the public skill ships only generic suites/examples.
Doctrine (from Anthropic — encode, don't restate)
- Grade the output/outcome, not the path. Tool-call-sequence asserts are brittle and demoted to opt-in; the everyday suite grades what the agent produced. The legacy
core-behaviors suite (tool-sequence graded) is retained only as an example of this anti-pattern — it is a v1 tasks: file and is not runnable by EvalRunner, which reports it as a named error rather than attempting it.
- Capability starts low (a hill to climb); regression targets ~100%; passing capability cases graduate into regression.
- pass^k for reliability, pass@k where one success suffices.
- Partial credit via assert weights. Balance should-do and should-not cases — one-sided evals create one-sided optimization.
- Judge discipline: distinct judge model, reason-then-score, forced structured verdict, an Unknown escape hatch.
- Never trust a score until you read transcripts — every run persists full case transcripts to
MEMORY/STATE/Evals-Results/<suite>/<run>/run.json.
Harness integration
- Config-change regression:
hooks/ConfigEvalFire.hook.ts → LIFEOS/TOOLS/ConfigEvalOnChange.ts fires the configured dispositions suite when a behaviour-defining file changes (default core-dispositions, the runnable v2 suite; override via LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/config.json config_change_suite — identity-bound suites live in that USER layer, never the public tree); regressions notify Pulse. Non-blocking, subscription-billed, debounced.
- ISA / Algorithm: an eval suite is the operational form of an ISA claim's falsifier (integration map kept on the maintainer machine — session notes, does not ship).
Legacy (v1, superseded)
The v1 grader-stack (Graders/, TrialRunner.ts) and the @langwatch/scenario path (ScenarioRunner.ts, LifeosAgentAdapter.ts, API-billed) predate the assertion-first rewrite. Prefer the v2 path above. The scenario path bills ANTHROPIC_API_KEY — do not use it for principal work.
Gotchas
- Single-shot agent-under-test narrates tool calls. Running the full agentic system prompt through tool-less inference makes the agent defer and simulate tool use instead of answering — which tanks "lead with the answer" style cases. EvalRunner injects an
[EVALUATION CONTEXT] no tools, answer directly suffix to fix this; keep it when authoring output-graded disposition cases.
judge_level must differ from agent_level (Anthropic: judge ≠ generator). Default agent=medium, judge=high.
- Unknown counts as a miss. A judge that can't verify an assertion returns UNKNOWN, scored as fail — conservative for regression, correct for gates.
- Deterministic asserts are free; use them first. Reserve model asserts (
llm-rubric/llm-assert) for nuance a code check can't capture.
is-json checks the whole output; contains-json checks for an embedded fragment. Don't use is-json on prose that merely mentions JSON.
Execution Log
After completing any workflow, append a single JSONL entry:
echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Evals","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
1---2name: evals3description: Assertion-first AI eval framework aligned to Anthropic's 'Demystifying evals for AI agents' — typed deterministic asserts + a forced-structured LLM judge over an input→assert case schema, pass^k/pass@k, capability vs regression suites, subscription-billed. USE WHEN eval, evaluate, benchmark, regression test, assertion, assert, llm-rubric, judge, pass@k, pass^k, grade output, compare prompts/models, test agent. NOT FOR scientific-method framing (use Science), property/mutation testing of code (use Hardening), or live UI verification (use Interceptor).4---5
6# Evals — Assertion-First AI Evaluation
7
8## What it is
9
10An eval gives an AI an input, then applies **assertions** to its output to measure success (Anthropic's definition). A case is `{id, prompt, assert:[...]}`. Each assertion is either **deterministic** (code, fast/free) or **model-graded** (an LLM judge). Cases run multiple trials; we report **pass^k** (all trials pass — the honest metric for a reliability-critical agent) and **pass@k** (any trial passes). Everything routes through `Inference.ts` — subscription-billed, no API-key path, no external deps.
11
12Grounded in Anthropic's current doctrine — [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents), [Define success criteria / develop tests](https://platform.claude.com/docs/en/docs/build-with-claude/develop-tests), and the `skill-creator` `{text, passed, evidence}` assertion convention. The typed-assert layer is promptfoo-shaped but our own TS.
13
14**Freshness contract:** "aligned to Anthropic's doctrine" is a live claim, not a snapshot. When designing a new suite class or touching the `## Doctrine` section below, re-fetch the Demystifying-evals doc and flag where it has moved past what's encoded here. Advisory only — report divergence, never auto-adopt, and an unreachable URL never blocks a run.
15
16## The canonical path (v2)
17
18| Tool | Role |
19|------|------|
20| `Tools/Assertions.ts` | Deterministic assert engine: `equals`, `contains`, `icontains`, `contains-all/any`, `regex`, `starts-with`, `ends-with`, `is-json`, `contains-json`, `max-length`, `min-length`, each with `not-` negation. Sync, no model call. |
21| `Tools/Judge.ts` | Model-graded asserts `llm-rubric` (1–5 → 0–1, threshold) and `llm-assert` (NL assertions → TRUE/FALSE/UNKNOWN). Forced-structured JSON verdict, reason-then-score, distinct judge level, **Unknown→miss** escape hatch. |
22| `Tools/EvalRunner.ts` | Loads a suite, runs the agent-under-test per case (single-shot inference against the target system prompt), applies asserts, computes pass^k/pass@k, persists transcripts + `latest.json`. |
23| `Tools/SuiteManager.ts` | Suite listing + saturation tracking. |
24| `Tools/FailureToTask.ts` | Convert real failures into cases (seed from 20–50 real failures). |
25
26```bash
27# Run a suite (USER-customization suites resolve before the skill's own)
28bun run ${LIFEOS_SKILL_DIR}/Tools/EvalRunner.ts -s <suite> [-t trials] [--json]
29# Sanity-check the assert engine / judge
30bun run ${LIFEOS_SKILL_DIR}/Tools/Assertions.ts # 16-case self-test
31bun run ${LIFEOS_SKILL_DIR}/Tools/Judge.ts # good-vs-bad discrimination
32```
33
34## Workflow Routing
35
36| Workflow | Trigger | File |
37|----------|---------|------|
38| **RunEval** | "run the eval", "run suite", "evaluate this", "grade output" | `Workflows/RunEval.md` |
39| **CreateUseCase** | "new eval", "create a suite", "eval for X", "what should I test" | `Workflows/CreateUseCase.md` |
40| **CreateJudge** | "write a judge", "llm-rubric", "grading criteria", "judge prompt" | `Workflows/CreateJudge.md` |
41| **ComparePrompts** | "compare prompts", "which prompt is better", "A/B this prompt" | `Workflows/ComparePrompts.md` |
42| **CompareModels** | "compare models", "which model is better", "is the cheaper rung enough" | `Workflows/CompareModels.md` |
43| **ViewResults** | "eval results", "how did it score", "show the last run", "saturation" | `Workflows/ViewResults.md` |
44| **CreateScenario** | "create a scenario", "multi-turn eval", "scenario test" | `Workflows/CreateScenario.md` |
45| **RunScenario** | "run the scenario", "run multi-turn" | `Workflows/RunScenario.md` |
46
47## Suite / case schema (assertion-first)
48
49```yaml
50name: my-suite
51type: regression # or capability
52pass_threshold: 0.75
53agent_level: medium # agent-under-test inference level
54judge_level: high # judge != generator (Anthropic best practice)
55trials: 3
56# system_prompt: optional override; default = live system prompt + DA identity
57cases:
58 - id: descriptive_name
59 prompt: "the user turn sent to the agent-under-test"
60 assert:
61 - type: not-contains # deterministic
62 value: "should work"
63 weight: 1
64 - type: llm-rubric # model-graded, weighted for partial credit
65 weight: 2
66 value: "Does the output tie any done-claim to verification evidence?"
67 - type: llm-assert
68 weight: 1
69 value: ["The output does not claim success without evidence"]
70 - id: should_not_case # balance: test should-do AND should-not
71 negative: true
72 prompt: "..."
73 assert: [...]
74```
75
76Identity-bound suites (e.g. {{DA_NAME}}'s dispositions) live in `LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/Suites/` — the public skill ships only generic suites/examples.
77
78## Doctrine (from Anthropic — encode, don't restate)
79
80- **Grade the output/outcome, not the path.** Tool-call-sequence asserts are brittle and demoted to opt-in; the everyday suite grades what the agent produced. The legacy `core-behaviors` suite (tool-sequence graded) is retained only as an example of this anti-pattern — it is a v1 `tasks:` file and is not runnable by `EvalRunner`, which reports it as a named error rather than attempting it.
81- **Capability starts low** (a hill to climb); **regression targets ~100%**; passing capability cases **graduate** into regression.
82- **pass^k for reliability**, pass@k where one success suffices.
83- **Partial credit** via assert weights. **Balance** should-do and should-not cases — one-sided evals create one-sided optimization.
84- **Judge discipline:** distinct judge model, reason-then-score, forced structured verdict, an **Unknown** escape hatch.
85- **Never trust a score until you read transcripts** — every run persists full case transcripts to `MEMORY/STATE/Evals-Results/<suite>/<run>/run.json`.
86
87## Harness integration
88
89- **Config-change regression:** `hooks/ConfigEvalFire.hook.ts` → `LIFEOS/TOOLS/ConfigEvalOnChange.ts` fires the configured dispositions suite when a behaviour-defining file changes (default `core-dispositions`, the runnable v2 suite; override via `LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/config.json` `config_change_suite` — identity-bound suites live in that USER layer, never the public tree); regressions notify Pulse. Non-blocking, subscription-billed, debounced.
90- **ISA / Algorithm:** an eval suite is the operational form of an ISA claim's falsifier (integration map kept on the maintainer machine — session notes, does not ship).
91
92## Legacy (v1, superseded)
93
94The v1 grader-stack (`Graders/`, `TrialRunner.ts`) and the `@langwatch/scenario` path (`ScenarioRunner.ts`, `LifeosAgentAdapter.ts`, API-billed) predate the assertion-first rewrite. Prefer the v2 path above. The scenario path bills `ANTHROPIC_API_KEY` — do not use it for principal work.
95
96## Gotchas
97
98- **Single-shot agent-under-test narrates tool calls.** Running the full agentic system prompt through tool-less inference makes the agent defer and simulate tool use instead of answering — which tanks "lead with the answer" style cases. EvalRunner injects an `[EVALUATION CONTEXT] no tools, answer directly` suffix to fix this; keep it when authoring output-graded disposition cases.
99- **`judge_level` must differ from `agent_level`** (Anthropic: judge ≠ generator). Default agent=medium, judge=high.
100- **Unknown counts as a miss.** A judge that can't verify an assertion returns UNKNOWN, scored as fail — conservative for regression, correct for gates.
101- **Deterministic asserts are free; use them first.** Reserve model asserts (`llm-rubric`/`llm-assert`) for nuance a code check can't capture.
102- **`is-json` checks the whole output; `contains-json` checks for an embedded fragment.** Don't use `is-json` on prose that merely mentions JSON.
103
104## Execution Log
105
106After completing any workflow, append a single JSONL entry:
107
108```bash
109echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Evals","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl
110```