AI Evals

Use when testing an LLM-backed feature, prompt, tool loop or multi-step agent, where the same input can produce different outputs and a prompt or model change can regress behaviour with no code diff. The behaviour spec that precedes the prompt, scenario datasets including adversarial and degradation classes, deterministic assertions over OpenTelemetry traces, calibrated LLM-as-judge, CI gates with baselines, human-in-the-loop, and the production scoring loop that turns incidents into scenarios.

konradcinkusz Updated

File contents

konradcinkusz/architecture-standards/tree/main/plugins/quality-and-process/skills/ai-evals commit 2c64e301b9

Frequently asked questions

npx skillmds@latest add konradcinkusz/ai-evals