LLM Eval Design

Build LLM evaluations from real failures with calibrated judges and regression gates that catch quality drift. Use when an LLM feature needs quality measurement or prompts change without anyone knowing what broke.

Amey-Thakur Updated

File contents

Amey-Thakur/AI-SKILLS/tree/main/skills/llm-engineering/llm-eval-design commit 997cb0b5ba

Frequently asked questions

npx skillmds@latest add amey-thakur/llm-eval-design