Mapo Mixed Advantage Policy Optimization

Dynamically reweight advantage functions based on trajectory certainty to improve policy optimization in foundation models. Addresses advantage reversion and mirror problems by mixing standardized and mean-normalized advantage formulations. Enables more stable gradient signals across high- and low-certainty samples.

adu2021 744b40f 14.1 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/mapo-mixed-advantage-policy-optimization commit 744b40f43d

Frequently asked questions

npx skillmds@latest add adu2021/mapo-mixed-advantage-policy-optimization