Rm Bench Eval

Evaluates reward models' ability to correctly identify preferred responses based on substantive content rather than superficial stylistic cues. It probes sensitivity to subtle correctness differences, resistance to verbosity/style bias, and performance across diverse domains like math, code, and safety. Use when the user wants to benchmark on RM-Bench, or asks about evaluating this task. Reports Average Accuracy.

qhjqhj00 7ad7dbb 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/rm-bench-eval commit 7ad7dbbe57

Frequently asked questions

npx skillmds add qhjqhj00/rm-bench-eval