Eval Harness

Measure recorded multi-agent run outputs with two deterministic offline checkers plus a judge seam. Structural eval scores observability-tracing JSONL on a 0-100 rubric with flags and a regression baseline. Grounding audit holds claim-with-citation outputs to their sources — every claim must cite a source that exists and substantiates it — catching uncited, mis-cited, and fabricated references. An LLM-judge seam is documented for content quality. Scores, flags and deltas are for humans and CI only — never for the agent being scored. Use when a pipeline is unmeasured, when you need a regression baseline, when a RAG/research agent must be held to its sources, or when a quality metric is being fed back to the agent it measures.

sharp-skills Updated

File contents

sharp-skills/skills/tree/main/skills/eval-harness commit a5d34b8c3e

Frequently asked questions

npx skillmds@latest add sharp-skills/eval-harness