Eval Harness

Use when you need to evaluate an LLM pipeline or AI feature systematically — sets up an eval harness with test cases, scoring rubrics, and pass/fail tracking rather than one-off manual spot-checks

drvoss c07dd58 18.0 KB Updated

File contents

drvoss/everything-copilot-cli/tree/main/skills/testing/eval-harness commit c07dd5853b

Frequently asked questions

npx skillmds@latest add drvoss/eval-harness