LLM As Judge Scorer

Design a reliable LLM-as-judge metric — a calibrated rubric, a clear scoring scale, and bias controls — and validate it against human labels before trusting it. Use when grading open-ended LLM output (summaries, answers, tone) that exact-match can't score.

imtiazrayhan Updated

File contents

imtiazrayhan/agentscamp-library/tree/main/skills/llm-as-judge-scorer commit 3523ba49b0

Frequently asked questions

npx skillmds@latest add imtiazrayhan/llm-as-judge-scorer