LLM Judge Auditor

Audit an LLM-as-judge setup for self-preference, position bias, and lack of human agreement before its scores are trusted. Use when a model grades model output, for pairwise preference evaluations, automated scoring, or when someone reports win rates from an AI judge. Refuses to accept judge scores with no measured human agreement.

ityaadiii Updated

File contents

ityaadiii/skills-that-say-i-dont-know/tree/main/skills/llm-judge-auditor commit c913683501

Frequently asked questions

npx skillmds@latest add ityaadiii/llm-judge-auditor