Eval Harness First

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

thedixitjain 0de6a76 3 files · 24.3 KB Updated 2 repo stars

File contents

thedixitjain/the-mega-skill-library/tree/main/library/ai-agents-and-harness/eval-harness-first commit 0de6a762c8

Frequently asked questions

npx skillmds add thedixitjain/eval-harness-first