LLM Eval Harness

Design an evaluation harness for an LLM-powered feature — a versioned golden set (representative + adversarial + regression cases), the cheapest adequate grading method per case, a metric with a pre-set pass bar and regression gate, and a failure taxonomy that targets iteration. Use when the user is building or tuning an LLM feature (prompt, RAG, agent, classifier) and needs evals, a way to test prompt/model changes, or to stop shipping quality regressions on vibes.

sananthanarayan 7f5bf5b 5 files · 17.8 KB Updated

File contents

sananthanarayan/skilldrop/tree/main/skills/llm-eval-harness commit 7f5bf5b59a

Frequently asked questions

npx skillmds@latest add sananthanarayan/llm-eval-harness