Justrl Simple Recipe

Demonstrate that simple single-stage RL with fixed hyperparameters matches complex multi-stage approaches for training small LLMs on mathematical reasoning. Use basic setup: GRPO algorithm, rule-based verification, 16K token context, standard training data without difficulty filtering. Train two 1.5B models to competitive performance using 2× less compute than sophisticated approaches.

adu2021 d82da62 3.5 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/justrl-simple-recipe commit d82da62d1f

Frequently asked questions

npx skillmds@latest add adu2021/justrl-simple-recipe