Rlvr Reasoning Eval

This evaluation probes the mathematical and out-of-domain reasoning capabilities of large language models trained with Reinforcement Learning with Verifiable Rewards (RLVR). It specifically tests how well entropy-aware credit assignment methods allocate learning signals across high-entropy tokens during chain-of-thought generation. Performance is measured by average accuracy and pass rate over multiple sampled reasoning paths. Use when the user wants to benchmark on AIME24, AIME25, AMC, MATH, Minerva, Olympiad, or asks about evaluating this task. Reports Avg@k, Pass@k.

qhjqhj00 f09bc9e 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/rlvr-reasoning-eval commit f09bc9edc7

Frequently asked questions

npx skillmds add qhjqhj00/rlvr-reasoning-eval