Slow Fast Policy Optimization Rl

Stabilize RL for LLM reasoning via three-phase decomposition: fast inner trajectory optimization, repositioning to manage off-policy drift, slow correction for stable updates. Achieve up to 2.80-point math reasoning gains over GRPO while reducing rollouts 4.93x and wall-clock time 4.19x via improved stability without changing reward structure.

adu2021 2cea39e 11.5 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/slow-fast-policy-optimization-rl commit 2cea39e4eb

Frequently asked questions

npx skillmds@latest add adu2021/slow-fast-policy-optimization-rl