LLM Evals Audit

Use this skill when a developer wants to check whether their existing evaluations are trustworthy and well-targeted. Triggers on: "audit my evals", "are my evals any good", "review my evaluation setup", "check my LLM judges", "are my evaluations reliable", "something feels off with my eval scores", "inherited an eval system", "my evals are passing but the product feels broken", "post-build eval check", "validate my eval pipeline". Inspects existing eval artifacts — judge prompts, annotation data, issue reports, alignment scores — and produces a prioritized findings report with a concrete fix for each problem. Do NOT use this to build new evals from scratch — use llm-evals-checklist first, then llm-judge-creator.

latitude-dev 2a80fc6 16.0 KB Updated

File contents

latitude-dev/eval-skills/tree/main/skills/llm-evals-audit commit 2a80fc672a

Frequently asked questions

npx skillmds@latest add latitude-dev/llm-evals-audit