Pairjudge Rm Eval

Evaluates whether a reward model can correctly select the right solution from a set of N generated candidates for mathematical reasoning problems. It probes the model's ability to perform pairwise correctness judgments and rank solutions without relying on arbitrary scalar scores. Use when the user wants to benchmark on MATH-500, Olympiad Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 db73bf0 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pairjudge-rm-eval commit db73bf02f1

Frequently asked questions

npx skillmds add qhjqhj00/pairjudge-rm-eval