Grpo Rlvr Training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

wshobson 913ef1d 3 files · 22.8 KB Updated

File contents

wshobson/agents commit 913ef1d44c

Frequently asked questions

npx skillmds add wshobson/grpo-rlvr-training