Ce Gppo Gradient Preserving Entropy Control

Control policy entropy dynamics in RL by reweighting gradients from clipped tokens. CE-GPPO preserves out-of-clip gradients with beta parameters to stabilize exploration-exploitation balance, preventing entropy collapse while maintaining training stability in LLM fine-tuning.

adu2021 745a81b 16.9 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/ce-gppo-gradient-preserving-entropy-control commit 745a81b484

Frequently asked questions

npx skillmds@latest add adu2021/ce-gppo-gradient-preserving-entropy-control