LLM Reasoning Rl Eval

This evaluation protocol assesses the mathematical reasoning and multi-step problem-solving capabilities of LLMs after reinforcement learning fine-tuning. It measures how well models generalize from a specialized arithmetic training task to standard academic and competitive math benchmarks. Use when the user wants to benchmark on GSM8K, BBH, MATH, MMLU-Pro, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2c84ebb 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llm-reasoning-rl-eval commit 2c84ebb9af

Frequently asked questions

npx skillmds add qhjqhj00/llm-reasoning-rl-eval