LLM Regression Runner

Use this skill when a developer wants to test a prompt change against a golden dataset and see what broke. Triggers on: "run my evals", "test this prompt change", "check for regressions", "did I break anything", "run regression tests", "test against golden dataset", "compare prompt versions", "is it safe to deploy", "run offline evals", "what changed after my prompt update", "eval before deploying". Runs a golden dataset against the current prompt, scores each case with available judges, compares results against a saved baseline, and produces a pass/fail report with a clear deploy recommendation.

latitude-dev b522432 10.1 KB Updated

File contents

latitude-dev/eval-skills/tree/main/skills/llm-regression-runner commit b5224327f1

Frequently asked questions

npx skillmds@latest add latitude-dev/llm-regression-runner