Judgebench Eval

Probes the factual and logical reliability of LLM-based judges and reward models by testing their ability to distinguish between objectively correct responses and subtly flawed ones across knowledge, reasoning, math, and coding domains. Use when the user wants to benchmark on JudgeBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8761eca 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/judgebench-eval commit 8761ecad95

Frequently asked questions

npx skillmds add qhjqhj00/judgebench-eval