Eval Harness

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. Reuses loopkit's verifier subagent as the grader — do not build a new one.

archive228 720386e 3.3 KB Updated

File contents

archive228/loopkit/tree/main/skills/eval-harness commit 720386ebcd

Frequently asked questions

npx skillmds@latest add archive228/eval-harness