Cx Calibration Agreement

Use to measure and diagnose grader agreement in support QA — human vs human, or an AI grader vs human reviewers — separating random disagreement from one grader being systematically harsher. Trigger for "run a calibration", "how do our reviewers compare", "is the AI grading too harshly", "our QA scores have too many false positives", "why do reviewers disagree", contested or overturned evaluations, or checking a new grader or model version.

rulebase-co Updated 1 repo stars

File contents

rulebase-co/rulebase-skills/tree/main/skills/quality-assurance/cx-calibration-agreement commit 67b9d72ee1

Frequently asked questions

npx skillmds@latest add rulebase-co/cx-calibration-agreement