Platform Note: State files use the platform's native directory:
.claude/(Claude Code),.codex/(Codex CLI), or.cursor/(Cursor IDE). Shared metrics + the eval journal live in.orchestrator/metrics/. Seeskills/_shared/platform-tools.md.
Eval Skill — Session-Process Evaluation (aiat-llm-eval/1.0)
On-demand, honest measurement of ONE completed orchestrator session against the
pre-registered rubric-v1 check set. The deterministic engine
(scripts/eval-session.mjs → scripts/lib/eval/engine.mjs) reads only local
metrics files (sessions.jsonl + events.jsonl), scores the five deterministic
dimensions, appends a session-eval record to the journal, and optionally
renders an HTML report. An opt-in LLM judge overlays two advisory dimensions.
The standard this skill implements is docs/eval/aiat-llm-eval-v1.md;
the frozen, content-hashed check set is skills/eval/rubric-v1.md.
Posture Contract (load-bearing — read before executing)
- No global score, by construction. The record has no overall/total/mean field, and this skill never derives one. Report per-dimension verdicts only.
- Never guess. Missing source data yields
cannot-determine(a first-class, non-error verdict) with an honest reason — never a fabricatedpass/fail. Do NOT "fill in" a missing KPI or infer a gate result the events do not show. - Deterministic before judge. The five deterministic dimensions are complete
on their own. The judge (Phase 3) is opt-in, ADVISORY, and
uncalibratedin v1 — never blend a judge verdict into the deterministic tally. - Journal is SSOT; the report is a derived view. The append-only
.orchestrator/metrics/eval.jsonlis authoritative. The HTML report is rebuildable from any stored record and is never authoritative over the journal. --verifyis the reproducibility proof. Re-scoring stored source data reproduces the stored dimensions byte-for-byte (exit 0) or reports drift (exit 1). This proves the SCORING replays — NOT that the model is deterministic.- Self-evaluation is labelled as such. The orchestrator scoring its own session is a self-evaluation, not an independent audit.
Phase 0: Bootstrap Gate
Read skills/_shared/bootstrap-gate.md and execute the gate check. If the gate is
CLOSED, invoke skills/bootstrap/SKILL.md and wait for completion before
proceeding. If the gate is OPEN, continue to Phase 1.
Phase 1: Config & Argument Loading
1.1 Read Session Config
Read and parse Session Config per skills/_shared/config-reading.md. Extract the
eval block (scripts/lib/config.mjs returns it as config.eval, parsed by
scripts/lib/config/eval.mjs):
enabled: boolean (default false)
mode: 'warn' | 'off' (default 'warn')
judge: 'off' | 'haiku' | 'sonnet' (default 'off')
report: 'html' | 'none' (default 'html')
handle: string | null (default null)
On-demand /eval runs regardless of eval.enabled. The enabled flag gates
the AUTOMATIC session-end eval phase only — it does NOT gate this command (same
posture as /reconcile vs reconcile.enabled). mode: off is honoured as a
kill-switch only for the automatic phase; on-demand invocation still runs. If
eval.judge is off, skip Phase 3 entirely.
Parser gotcha: the
eval:key-line itself MUST NOT carry an inline comment (strict/^eval:\s*$/); a trailing# commenton that exact line makes the parser skip the whole block and silently apply ALL defaults. Sub-key lines tolerate inline comments.
1.2 Parse Arguments
Inspect $ARGUMENTS:
--session <id>→ pass through to--session.--no-write→ evaluate without appending to the journal (dry-run).--verify <run-id>→ verification mode: skip Phases 2–4, run the CLI--verifypath (see Phase 6), report MATCH/DRIFT, done.
1.3 Capture the Model Id (honest provenance)
The record's model.source records HOW the model id was captured, precisely
because self-report is unreliable:
- If
$ANTHROPIC_MODELis set in the environment, the engine reads it automatically withsource: env— env wins over the flag (precedenceenv > flag). Do not pass--model-idin that case; let the engine resolve it. - Otherwise the coordinator passes its own self-reported model id:
--model-id <self-reported-model-id> --model-source self-report.
Phase 2: Deterministic Run
Run the deterministic engine via its CLI. Default target is the last completed
session (resolution cascade); --session overrides.
node scripts/eval-session.mjs [--session <id>] --json \
[--model-id <self-reported-id> --model-source self-report] \
[--no-write]
- Do NOT pass
--metrics-dirfor a real run — the engine defaults to the live.orchestrator/metrics, the session being evaluated. - The CLI captures the eval
timestamp(the one sanctioned clock read) and hands it to the engine as a parameter, so the scoring path stays clock-free and--verify-reproducible. - Exit codes:
0success ·1user error (session not found) ·2system error. On exit1(e.g. "no completed session found"), surface the message and stop — do not retry with fabricated inputs.
Parse the emitted JSON record. It carries dimensions[] (5 deterministic
entries), kpis{}, provenance.rubric_sha256 (non-null once rubric-v1.md
exists), model, harness, and run_id. Unless --no-write was passed, the
record is already appended to .orchestrator/metrics/eval.jsonl by the CLI.
Contamination check: if the human-render/summary reports a peer-overlapped
window, note it — verification-evidence and gate-health will read
cannot-determine for that reason (attribution is unsafe), which is correct, not
a defect.
Phase 3: Judge Overlay (ONLY when eval.judge != off)
The judge runs coordinator-side — AskUserQuestion and the Agent tool are
not available inside a dispatched subagent, so the judge is dispatched from the
coordinator thread using the read-only agent session-orchestrator:eval-judge
(model = eval.judge). Reference the API; do not reimplement scoring here:
import { runEvalJudge, mergeJudgeDimensions } from '$PLUGIN_ROOT/scripts/lib/eval/judge.mjs';
import { appendEvalRecord } from '$PLUGIN_ROOT/scripts/lib/eval/sink.mjs';
// dispatchAgent = the coordinator's Agent-tool dispatch closure targeting
// subagent_type 'session-orchestrator:eval-judge'.
const { status, dimensions } = await runEvalJudge({
dispatchAgent,
record, // the deterministic record from Phase 2
model: EVAL_JUDGE_MODEL, // eval.judge ('haiku' | 'sonnet')
budget: JUDGE_BUDGET_TOKENS, // optional
});
// mergeJudgeDimensions appends the advisory judge dimensions to the record.
const merged = mergeJudgeDimensions(record, dimensions);
// The COORDINATOR appends the enriched record (subagents never write the journal).
appendEvalRecord(merged, { path: '.orchestrator/metrics/eval.jsonl' });
- Every judge dimension arrives
advisory: true+calibration_status: "uncalibrated"(the schema firewall rejects any other shape). Keep them visibly separated from the deterministic five in the summary. - If
runEvalJudgereturns a non-okstatus(e.g. dispatch failed), keep the deterministic record as-is and note the judge was unavailable — the deterministic evaluation is complete without it. - When
--no-writewas passed in Phase 2, do NOT append the merged record either.
The judge merge re-writes the record with the SAME
run_id/timestamp, so a later--verify <run-id>re-scores the deterministic dimensions from source and diffs them; judge dimensions are advisory and excluded from the drift diff.
Phase 4: Report (ONLY when eval.report == html)
Render the derived HTML view from the record:
import { writeEvalReport } from '$PLUGIN_ROOT/scripts/lib/eval/report.mjs';
const res = writeEvalReport(record, { generatedAt: new Date().toISOString() });
// res.ok === true → res.path === .orchestrator/eval/reports/<run_id>.html
- Output path:
.orchestrator/eval/reports/<run_id>.html(gitignored — a derived view, rebuildable from the journal). writeEvalReportNEVER throws; onres.ok === falsesurface the WARN reason and continue (the journal record is unaffected — the report is derived).- Name the report path in the chat so the operator can open it.
- When
eval.report == none, skip this phase.
Phase 5: Chat Summary
Emit a compact, honest per-dimension summary. Status lines only — no global score.
## /eval — <session_id> (self-evaluation, aiat-llm-eval/1.0 · rubric-v1 · n=1, no CI)
Deterministic:
verification-evidence PASS <one-line evidence>
plan-fidelity PASS completion_rate=1.0 (score)
gate-health PASS <one-line evidence>
process-safety PASS <one-line evidence + guard-emission disclosure>
efficiency-kpis N/A (reported: duration=…s waves=… agents=… tok_in=… tok_out=… carryover=…)
Judge (advisory, uncalibrated) [only when eval.judge != off]:
instruction-adherence <verdict> advisory
report-quality <verdict> advisory
cannot-determine: <k> of 5 deterministic dimensions (<reasons>)
Report: .orchestrator/eval/reports/<run_id>.html
Journal: .orchestrator/metrics/eval.jsonl (appended: <yes|--no-write>)
Re-verify: node scripts/eval-session.mjs --verify <run_id>
Always report the cannot-determine share explicitly — a high abstention count
is an honest signal about missing telemetry, not a failure to hide. Always print
the --verify command as the reproducibility handle.
Phase 6: Verification Mode (--verify <run-id>)
When Phase 1.2 detected --verify, run ONLY:
node scripts/eval-session.mjs --verify <run-id> --json
- Exit
0+{ match: true }→ the stored record re-scores identically across all deterministic dimensions. Report MATCH with the dimension count. - Exit
1+{ match: false, diffs }→ scoring drift. Report the per-dimension diff (id.field: stored=… fresh=…). Drift means the source data or the engine changed since the record was written — investigate, do not overwrite. --verifyreproduces the stored model + timestamp verbatim (no env override), so a MATCH is a real reproducibility proof of the scoring, not of model output.
Cross-Platform (Codex CLI / Cursor / Pi) — FA4
The deterministic core is pure Node CLIs (scripts/eval-session.mjs) plus Node
library modules (report.mjs, sink.mjs) — they run identically on every
platform. Only the judge phase needs harness-specific tooling.
- Codex CLI / Cursor / Pi: the
Agenttool (andAskUserQuestion) are unavailable, so Phase 3 (judge) is SKIPPED with a one-line note ("judge phase skipped: requires the Agent tool, unavailable on<platform>"). Phases 2, 4, 5, 6 run unchanged — they are Node-only. Seeskills/_shared/platform-tools.md§ Agent Dispatch Pattern. harness.platformon the record is resolved from$SO_PLATFORM(falls back toclaude-code) inside the engine — no skill action needed.- The deterministic five dimensions + the HTML report +
--verifyare fully available on all platforms; the judge overlay is a Claude-Code-only enrichment in v1.
Anti-Patterns
- DO NOT derive, print, or imply a global/overall/aggregate score — the record forbids one by construction and so does every report.
- DO NOT guess or "fill in" missing data — a missing gate result, KPI, or
completion_rate is
cannot-determine/null, never a fabricated pass or0. - DO NOT present the judge's advisory verdict as a measurement — it is
uncalibratedin v1 and must stay visibly separated from the deterministic tally. - DO NOT treat the HTML report as authoritative — the journal is the SSOT; the report is a rebuildable derived view.
- DO NOT skip
--verifywhen reproducibility is in question — it is the executable proof, and its MATCH/DRIFT exit code is the source of truth. - DO NOT pass
--metrics-dirfor a real run — that points the engine at fixture data instead of the live session metrics. - DO NOT call
runReconcile-style writes from a subagent — the coordinator owns everyeval.jsonlappend (PSA-007).
Source: hashgraph-online/awesome-codex-plugins → plugins/Kanevry/session-orchestrator/skills/eval/SKILL.md