Judge Calibration Auditor

Iterate-stage skill: analyzes human labels vs LLM-judge verdicts and turns every disagreement into a classified calibration signal with a stated correction — never auto-resolved toward the judge. Use when a judge and humans diverge — 'our judge disagrees with human reviewers', 'human labels vs judge verdicts, what's drifting', 'the judge scores everything 4', 'disagreement analysis' — or when /pm routes such a request here. Do NOT use to design the judge (llm-as-judge-designer), to build the eval (eval-engine), to produce the human labels themselves, or for judge-bias knowledge questions.

Abhillashjadhav Updated

File contents

Abhillashjadhav/PM-agent-OS/tree/main/.claude/skills/judge-calibration-auditor commit 0abefccc03

Frequently asked questions

npx skillmds@latest add abhillashjadhav/judge-calibration-auditor