LLM As Judge

Use an LLM as an evaluator for open-ended outputs — rubrics, pairwise comparison, calibration with human labels, bias mitigation. Covers when LLM-judge works, when it fails, and how to trust its scores. Use this skill when evaluating generative outputs at scale, building eval pipelines, or replacing expensive human review for non-critical judgments. Activate when: LLM as judge, LLM evaluator, automated evaluation, pairwise comparison, rubric evaluation, eval model.

latestaiagents Updated

File contents

latestaiagents/agent-skills/tree/main/skills/evals/llm-as-judge commit 060c66b943

Frequently asked questions

npx skillmds@latest add latestaiagents/llm-as-judge