Evaluation

Use when scoring a completed agent task, implementation, document, skill upgrade, or other deliverable against the original request, acceptance criteria, verification evidence, quality rubric, and residual risks before calling it done. Covers skeptical critic review, 1-5 scoring, score ceilings, evidence sufficiency, finding/action capture, and the evaluation-revision loop. Do NOT use for designing eval datasets or graders (use eval-driven-development), line-by-line diff review (use code-review), choosing test levels (use testing-strategy), or designing the overall process and gates before work starts (use methodology).

jacob-balslev Updated

File contents

jacob-balslev/skills/tree/main/skills/ai-engineering/evaluation commit 1c4dad868f

Frequently asked questions

npx skillmds@latest add jacob-balslev/evaluation-2