LLM Eval Harness

Measures how far a cheaper model falls from the Fable-5 quality bar on real task types — deterministic CHECKS (format, discipline, no-fabrication, surgical scope) plus line-similarity to captured Fable goldens; appends every run to a ratchet so the gap is trackable over time. No API key, no LLM-judge. Use when: "eval the model", "measure the Fable gap", "run the eval harness", "score this output", or deciding whether a cheaper model is good enough to switch to.

Evan-Daruwalla 0c75437 5 files · 25.9 KB Updated

File contents

Evan-Daruwalla/claude-skill-suite/tree/main/llm-eval-harness commit 0c75437559

Frequently asked questions

npx skillmds@latest add evan-daruwalla/llm-eval-harness