Outcome Eval
Post-execution LLM-judgment: did the implementation actually satisfy its spec? Reads the spec's acceptance section, the change diff, and test output, and emits a confidence-rated OutcomeVerdict (SATISFIED | NOT_SATISFIED | INCONCLUSIVE) with a rationale and unmet criteria. Ship authority is derived in TypeScript, never trusted from the LLM: a high-confidence NOT_SATISFIED blocks ship; every other verdict is advisory. The harness's first blocking post-execution spec-satisfaction gate (long a top-priority gap). Each verdict persists as an execution_outcome node, compounding into skill-effectiveness baselines.
When to Use
- At orchestrator step 6.5 — after Code Review, before Ship — on every change with a spec.
- When you need a durable, structured answer to "did this code do what the spec said?"
- NOT for pre-execution risk simulation (use PESL).
- NOT for rule-based floors (lint/architecture/entropy) or craft ceilings (naming/spec/security) — those run elsewhere.
- NOT for auto-remediation. outcome-eval judges; it does not fix.
- NOT when no judgable spec section exists — the verdict degrades to INCONCLUSIVE/advisory and never blocks.
Process
Phase 1: GATHER — Collect inputs
- Capture the change under judgment as a unified diff:
git diff (or git diff <base>...HEAD for a branch). Record it as diff.
- Capture test-runner output. If a test command is known, run it and capture stdout+stderr as
testOutput; otherwise pass the most recent captured output. Empty/unparseable test output is tolerated (degrades to advisory).
- Resolve the spec path. Prefer the spec under
docs/changes/<feature>/proposal.md for the current change. Record as specPath.
Phase 2: RESOLVE — Find the judgment section
The evaluator resolves the section internally via the fallback chain ## Success Criteria -> ## User-Visible Behavior -> ## Overview, recording the match in judgedAgainst. No manual action — pass specPath and let OutcomeEvaluator resolve. If no section is judgable, the verdict is INCONCLUSIVE/advisory.
Phase 3: JUDGE — Invoke the evaluator
- Invoke the MCP tool
mcp__harness__outcome_eval with { specPath, diff, testOutput } (optional model). The tool constructs OutcomeEvaluator cli-side and calls evaluate({ specPath, diff, testOutput }); the supported v1 provider is the anthropic analysis provider (ANTHROPIC_API_KEY).
diff and testOutput are required inputs and the agent MUST supply them from the session (git diff + captured test-runner output). They are the evidence the judge reasons over — passing an empty diff or empty testOutput is the degradation path, not the normal path: the verdict degrades to INCONCLUSIVE/advisory (never blocking), which defeats the gate. Do not invoke the tool without real diff/test content.
- The LLM returns ONLY
verdict / confidence / rationale / unmetCriteria. authority is computed in TypeScript from (verdict, confidence) and is never read from the LLM — do not attempt to override it. The tool returns the verdict exactly as the evaluator derives it.
- The call is degrade-safe: provider failure (incl. no
ANTHROPIC_API_KEY), empty diff, empty test output, or missing judgable section yields INCONCLUSIVE/low/advisory. It never throws and never blocks.
Phase 4: GATE — Render and (conditionally) halt
- Render the verdict:
verdict, confidence, judgedAgainst, rationale, and unmetCriteria.
- Authority rule (must match
deriveAuthority): authority is blocking iff verdict === 'NOT_SATISFIED' && confidence === 'high'; every other combination — including all INCONCLUSIVE and SATISFIED cases, and all medium/low NOT_SATISFIED — is advisory.
- On a blocking verdict: HALT before the Ship step. Report the unmet criteria and stop; do not proceed to step 7. Resolution requires fixing the implementation (or the spec) and re-running outcome-eval.
- On an advisory verdict: report it and proceed. Advisory
NOT_SATISFIED is surfaced for human attention but does not stop the workflow.
Harness Integration
mcp__harness__outcome_eval — MCP tool (the invocation surface). Inputs: specPath (required), diff (required), testOutput (required), model (optional), path (optional project root for graph persistence). The agent supplies diff and testOutput from the session; omitting them degrades the verdict to INCONCLUSIVE/advisory (never blocking). The handler builds the cli AnalysisProvider + a GraphStore, constructs OutcomeEvaluator, and returns the OutcomeVerdict with authority exactly as derived in TypeScript.
- Evaluator surface:
OutcomeEvaluator, deriveAuthority, verdictSchema, OutcomeVerdict are exported from @harness-engineering/intelligence.
- Provider path (v1 supported): the anthropic analysis provider (
ANTHROPIC_API_KEY). When no provider is configured the call degrades to INCONCLUSIVE/advisory. The openai-compatible strict structured-output path is a known follow-up (see Known Limitations).
- Orchestrator: runs as step 6.5 between Code Review and Ship in
harness.orchestrator.md.
- Persistence: each
evaluate() writes one execution_outcome node via ExecutionOutcomeConnector, consumable by effectiveness/scorer.ts.
- Relationship to
acceptance-eval: acceptance-eval is the upstream twin — it gates spec measurability before execution; outcome-eval judges implementation satisfaction after. The "authority is never read from the LLM" discipline spans both.
Known Limitations
- INCONCLUSIVE persistence: the persisted node maps
INCONCLUSIVE -> result: 'failure' for type-validity, but it OMITS agentPersona and writes affectedSystemNodeIds: []. The effectiveness scorer (gatherOutcomes) ignores any node missing agentPersona or outcome_of edges, so outcome-eval nodes are scorer-non-counting in v1 — the INCONCLUSIVE-as-failure mapping is therefore harmless and does not punish any persona. If a future change attaches persona/affected-system attribution, it MUST first change INCONCLUSIVE modeling (do not persist INCONCLUSIVE, or use a distinct result value the scorer excludes) before the node becomes scorer-counted.
- openai-compatible strict mode:
zodToJsonSchema does not emit additionalProperties: false, which OpenAI strict structured output requires. The v1 supported path is claude-cli / anthropic. Follow-up tracked.
- CI required-check wiring: deferred to a future CI workflow template.
Success Criteria
See docs/changes/outcome-eval/proposal.md for the full 9 criteria. This skill satisfies SC8 (orchestrator step 6.5 + blocking halt) and SC9 (introduces no new harness validate findings; layer rules respected).
Rationalizations to Reject
These are common rationalizations that sound reasonable but lead to incorrect results. When you catch yourself thinking any of these, stop and follow the documented process instead.
| Rationalization |
Why It Is Wrong |
"I don't have test output on hand, so I'll invoke the tool with just the diff." |
diff AND testOutput are both required evidence. Passing an empty testOutput is the degradation path — the verdict silently falls to INCONCLUSIVE/advisory, a false-negative at the ship gate. Gather real test output from the session before invoking. |
| "The verdict is high-confidence NOT_SATISFIED, but the author insists it's fine, so I'll downgrade it to advisory and proceed to Ship." |
authority is deriveAuthority(verdict, confidence) in TypeScript, never read from the LLM or the author. A high-confidence NOT_SATISFIED blocks. The only resolution is to fix the implementation (or amend the spec on its own merits) and re-run — not to override the gate. |
| "This is a small change and Code Review already passed, so I can skip step 6.5." |
The gate runs at orchestrator step 6.5, after Code Review and before Ship, and is not optional. A blocking verdict halts before step 7 regardless of change size. |
| "The model keeps returning INCONCLUSIVE, so I'll re-prompt until it commits to SATISFIED." |
Repeated INCONCLUSIVE almost always means diff/testOutput weren't supplied or the spec has no judgable section — not that the prompt is too strict. Fix the inputs; never loosen the conservative-confidence prompt to force a verdict. |
"INCONCLUSIVE persists as result: 'failure', so I'll suppress persistence to avoid punishing the persona." |
outcome-eval nodes are scorer-non-counting in v1 (they omit agentPersona and outcome_of edges), so the mapping punishes no one. Do not alter persistence modeling ad hoc; that must change only alongside a deliberate scorer-attribution change. |
Examples
Example: NOT_SATISFIED with high confidence (blocks)
Input: spec Success Criteria require GET /api/users/:id to return 404 with { error: 'User not found' }; the diff implements the happy path only, no 404 branch; test output shows the 404 test failing.
Verdict:
verdict: NOT_SATISFIED
confidence: high
judgedAgainst: success-criteria
authority: blocking
unmetCriteria:
- "404 path for nonexistent user is unimplemented; the failing test asserts { error: 'User not found' }."
rationale: "The diff adds the lookup but returns 200 with an empty body when the user is missing."
Action: HALT before Ship. Report unmet criteria; do not open the PR.
Example: partial implementation (advisory)
Input: the diff meets most criteria; one acceptance item is ambiguous in the diff.
Verdict: NOT_SATISFIED confidence: medium authority: advisory — surfaced for review, workflow proceeds.
Gates
- Authority is never read from the LLM. The verdict's
authority is always deriveAuthority(verdict, confidence) computed in TypeScript. If you find yourself letting the model assert blocking/advisory, STOP — that defeats the entire purpose of this gate.
- Block only on high-confidence NOT_SATISFIED.
authority === 'blocking' iff verdict === 'NOT_SATISFIED' && confidence === 'high'. Every other combination — all SATISFIED, all INCONCLUSIVE, and every medium/low NOT_SATISFIED — is advisory. Do not halt the workflow on an advisory verdict.
- Always supply
diff and testOutput. Omitting them degrades the verdict to INCONCLUSIVE/advisory (a silent false-negative at the ship gate). Gather them from the session before invoking the tool.
- Never block on infrastructure noise. A provider failure, an unparseable response, or a missing spec must resolve to INCONCLUSIVE/advisory, never a thrown error or a block. The evaluator enforces this; do not reintroduce a hard failure in the wrapper.
- Do not skip step 6.5. The gate runs after Code Review and before Ship. A blocking verdict halts before step 7; it is not optional.
Escalation
- Blocking verdict the author disputes: the resolution is to fix the implementation (or amend the spec's Success Criteria if they were wrong) and re-run outcome-eval — not to override the gate. If the spec itself is wrong, that is a spec change, reviewed on its own merits.
- Repeated INCONCLUSIVE on a real change: usually means
diff/testOutput were not supplied, or no judgable section exists in the spec. Confirm inputs and that the spec has a Success Criteria / User-Visible Behavior / Overview section.
- No
ANTHROPIC_API_KEY configured: every verdict degrades to INCONCLUSIVE/advisory and nothing blocks. Surface this to the human — the gate is effectively disabled until a provider is configured.
- Verdict seems wrong (false positive/negative): capture the spec section, diff, and verdict, and route to the maintainers; do not loosen the conservative-confidence prompt ad hoc.
1---2name: outcome-eval3description: Outcome Eval4---5# Outcome Eval67> Post-execution LLM-judgment: did the implementation actually satisfy its spec? Reads the spec's acceptance section, the change diff, and test output, and emits a confidence-rated `OutcomeVerdict` (`SATISFIED | NOT_SATISFIED | INCONCLUSIVE`) with a rationale and unmet criteria. Ship authority is derived in TypeScript, never trusted from the LLM: a high-confidence `NOT_SATISFIED` blocks ship; every other verdict is advisory. The harness's first blocking post-execution spec-satisfaction gate (long a top-priority gap). Each verdict persists as an `execution_outcome` node, compounding into skill-effectiveness baselines.89## When to Use1011- At orchestrator step 6.5 — after Code Review, before Ship — on every change with a spec.12- When you need a durable, structured answer to "did this code do what the spec said?"13- NOT for pre-execution risk simulation (use PESL).14- NOT for rule-based floors (lint/architecture/entropy) or craft ceilings (naming/spec/security) — those run elsewhere.15- NOT for auto-remediation. outcome-eval judges; it does not fix.16- NOT when no judgable spec section exists — the verdict degrades to INCONCLUSIVE/advisory and never blocks.1718## Process1920### Phase 1: GATHER — Collect inputs21221. Capture the change under judgment as a unified diff: `git diff` (or `git diff <base>...HEAD` for a branch). Record it as `diff`.232. Capture test-runner output. If a test command is known, run it and capture stdout+stderr as `testOutput`; otherwise pass the most recent captured output. Empty/unparseable test output is tolerated (degrades to advisory).243. Resolve the spec path. Prefer the spec under `docs/changes/<feature>/proposal.md` for the current change. Record as `specPath`.2526### Phase 2: RESOLVE — Find the judgment section2728The evaluator resolves the section internally via the fallback chain `## Success Criteria` -> `## User-Visible Behavior` -> `## Overview`, recording the match in `judgedAgainst`. No manual action — pass `specPath` and let `OutcomeEvaluator` resolve. If no section is judgable, the verdict is INCONCLUSIVE/advisory.2930### Phase 3: JUDGE — Invoke the evaluator31321. Invoke the MCP tool `mcp__harness__outcome_eval` with `{ specPath, diff, testOutput }` (optional `model`). The tool constructs `OutcomeEvaluator` cli-side and calls `evaluate({ specPath, diff, testOutput })`; the supported v1 provider is the anthropic analysis provider (`ANTHROPIC_API_KEY`).332. **`diff` and `testOutput` are required inputs and the agent MUST supply them** from the session (`git diff` + captured test-runner output). They are the evidence the judge reasons over — passing an empty `diff` or empty `testOutput` is the degradation path, not the normal path: the verdict degrades to INCONCLUSIVE/advisory (never blocking), which defeats the gate. Do not invoke the tool without real diff/test content.343. The LLM returns ONLY `verdict / confidence / rationale / unmetCriteria`. `authority` is computed in TypeScript from `(verdict, confidence)` and is never read from the LLM — do not attempt to override it. The tool returns the verdict exactly as the evaluator derives it.354. The call is degrade-safe: provider failure (incl. no `ANTHROPIC_API_KEY`), empty diff, empty test output, or missing judgable section yields INCONCLUSIVE/low/advisory. It never throws and never blocks.3637### Phase 4: GATE — Render and (conditionally) halt38391. Render the verdict: `verdict`, `confidence`, `judgedAgainst`, `rationale`, and `unmetCriteria`.402. Authority rule (must match `deriveAuthority`): authority is `blocking` **iff** `verdict === 'NOT_SATISFIED' && confidence === 'high'`; every other combination — including all `INCONCLUSIVE` and `SATISFIED` cases, and all `medium`/`low` `NOT_SATISFIED` — is `advisory`.413. **On a blocking verdict: HALT before the Ship step.** Report the unmet criteria and stop; do not proceed to step 7. Resolution requires fixing the implementation (or the spec) and re-running outcome-eval.424. On an advisory verdict: report it and proceed. Advisory `NOT_SATISFIED` is surfaced for human attention but does not stop the workflow.4344## Harness Integration4546- **`mcp__harness__outcome_eval`** — MCP tool (the invocation surface). Inputs: `specPath` (required), `diff` (required), `testOutput` (required), `model` (optional), `path` (optional project root for graph persistence). The agent supplies `diff` and `testOutput` from the session; omitting them degrades the verdict to INCONCLUSIVE/advisory (never blocking). The handler builds the cli `AnalysisProvider` + a `GraphStore`, constructs `OutcomeEvaluator`, and returns the `OutcomeVerdict` with authority exactly as derived in TypeScript.47- **Evaluator surface:** `OutcomeEvaluator`, `deriveAuthority`, `verdictSchema`, `OutcomeVerdict` are exported from `@harness-engineering/intelligence`.48- **Provider path (v1 supported):** the anthropic analysis provider (`ANTHROPIC_API_KEY`). When no provider is configured the call degrades to INCONCLUSIVE/advisory. The openai-compatible _strict_ structured-output path is a known follow-up (see Known Limitations).49- **Orchestrator:** runs as step 6.5 between Code Review and Ship in `harness.orchestrator.md`.50- **Persistence:** each `evaluate()` writes one `execution_outcome` node via `ExecutionOutcomeConnector`, consumable by `effectiveness/scorer.ts`.51- **Relationship to `acceptance-eval`:** `acceptance-eval` is the upstream twin — it gates spec measurability before execution; `outcome-eval` judges implementation satisfaction after. The "authority is never read from the LLM" discipline spans both.5253## Known Limitations5455- **INCONCLUSIVE persistence:** the persisted node maps `INCONCLUSIVE -> result: 'failure'` for type-validity, but it OMITS `agentPersona` and writes `affectedSystemNodeIds: []`. The effectiveness scorer (`gatherOutcomes`) ignores any node missing `agentPersona` or `outcome_of` edges, so outcome-eval nodes are **scorer-non-counting** in v1 — the INCONCLUSIVE-as-failure mapping is therefore harmless and does not punish any persona. If a future change attaches persona/affected-system attribution, it MUST first change INCONCLUSIVE modeling (do not persist INCONCLUSIVE, or use a distinct result value the scorer excludes) before the node becomes scorer-counted.56- **openai-compatible strict mode:** `zodToJsonSchema` does not emit `additionalProperties: false`, which OpenAI strict structured output requires. The v1 supported path is claude-cli / anthropic. Follow-up tracked.57- **CI required-check wiring:** deferred to a future CI workflow template.5859## Success Criteria6061See `docs/changes/outcome-eval/proposal.md` for the full 9 criteria. This skill satisfies SC8 (orchestrator step 6.5 + blocking halt) and SC9 (introduces no new `harness validate` findings; layer rules respected).6263## Rationalizations to Reject6465These are common rationalizations that sound reasonable but lead to incorrect results. When you catch yourself thinking any of these, stop and follow the documented process instead.6667| Rationalization | Why It Is Wrong |68| --------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |69| "I don't have test output on hand, so I'll invoke the tool with just the `diff`." | `diff` AND `testOutput` are both required evidence. Passing an empty `testOutput` is the degradation path — the verdict silently falls to INCONCLUSIVE/advisory, a false-negative at the ship gate. Gather real test output from the session before invoking. |70| "The verdict is high-confidence NOT_SATISFIED, but the author insists it's fine, so I'll downgrade it to advisory and proceed to Ship." | `authority` is `deriveAuthority(verdict, confidence)` in TypeScript, never read from the LLM or the author. A high-confidence NOT_SATISFIED blocks. The only resolution is to fix the implementation (or amend the spec on its own merits) and re-run — not to override the gate. |71| "This is a small change and Code Review already passed, so I can skip step 6.5." | The gate runs at orchestrator step 6.5, after Code Review and before Ship, and is not optional. A blocking verdict halts before step 7 regardless of change size. |72| "The model keeps returning INCONCLUSIVE, so I'll re-prompt until it commits to SATISFIED." | Repeated INCONCLUSIVE almost always means `diff`/`testOutput` weren't supplied or the spec has no judgable section — not that the prompt is too strict. Fix the inputs; never loosen the conservative-confidence prompt to force a verdict. |73| "INCONCLUSIVE persists as `result: 'failure'`, so I'll suppress persistence to avoid punishing the persona." | outcome-eval nodes are scorer-non-counting in v1 (they omit `agentPersona` and `outcome_of` edges), so the mapping punishes no one. Do not alter persistence modeling ad hoc; that must change only alongside a deliberate scorer-attribution change. |7475## Examples7677### Example: NOT_SATISFIED with high confidence (blocks)7879**Input:** spec Success Criteria require `GET /api/users/:id` to return 404 with `{ error: 'User not found' }`; the diff implements the happy path only, no 404 branch; test output shows the 404 test failing.8081**Verdict:**8283```84verdict: NOT_SATISFIED85confidence: high86judgedAgainst: success-criteria87authority: blocking88unmetCriteria:89 - "404 path for nonexistent user is unimplemented; the failing test asserts { error: 'User not found' }."90rationale: "The diff adds the lookup but returns 200 with an empty body when the user is missing."91```9293**Action:** HALT before Ship. Report unmet criteria; do not open the PR.9495### Example: partial implementation (advisory)9697**Input:** the diff meets most criteria; one acceptance item is ambiguous in the diff.9899**Verdict:** `NOT_SATISFIED confidence: medium authority: advisory` — surfaced for review, workflow proceeds.100101## Gates102103- **Authority is never read from the LLM.** The verdict's `authority` is always `deriveAuthority(verdict, confidence)` computed in TypeScript. If you find yourself letting the model assert blocking/advisory, STOP — that defeats the entire purpose of this gate.104- **Block only on high-confidence NOT_SATISFIED.** `authority === 'blocking'` iff `verdict === 'NOT_SATISFIED' && confidence === 'high'`. Every other combination — all `SATISFIED`, all `INCONCLUSIVE`, and every `medium`/`low` `NOT_SATISFIED` — is advisory. Do not halt the workflow on an advisory verdict.105- **Always supply `diff` and `testOutput`.** Omitting them degrades the verdict to INCONCLUSIVE/advisory (a silent false-negative at the ship gate). Gather them from the session before invoking the tool.106- **Never block on infrastructure noise.** A provider failure, an unparseable response, or a missing spec must resolve to INCONCLUSIVE/advisory, never a thrown error or a block. The evaluator enforces this; do not reintroduce a hard failure in the wrapper.107- **Do not skip step 6.5.** The gate runs after Code Review and before Ship. A blocking verdict halts before step 7; it is not optional.108109## Escalation110111- **Blocking verdict the author disputes:** the resolution is to fix the implementation (or amend the spec's Success Criteria if they were wrong) and re-run outcome-eval — not to override the gate. If the spec itself is wrong, that is a spec change, reviewed on its own merits.112- **Repeated INCONCLUSIVE on a real change:** usually means `diff`/`testOutput` were not supplied, or no judgable section exists in the spec. Confirm inputs and that the spec has a Success Criteria / User-Visible Behavior / Overview section.113- **No `ANTHROPIC_API_KEY` configured:** every verdict degrades to INCONCLUSIVE/advisory and nothing blocks. Surface this to the human — the gate is effectively disabled until a provider is configured.114- **Verdict seems wrong (false positive/negative):** capture the spec section, diff, and verdict, and route to the maintainers; do not loosen the conservative-confidence prompt ad hoc.