Pepo Token Level Multimodal Policy

Replace uniform token-level advantages with perception-exploration gating that weights tokens by visual grounding strength. Adds 3.67 points to geometry reasoning and 5.32 to few-shot classification with <1% compute overhead. Works best for multimodal CoT where visual grounding anchors reasoning. Trigger: When using token-level RL on VLMs and want to emphasize visually-grounded reasoning steps.

adu2021 Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.3-claude-opus-4.6/pepo-token-level-multimodal-policy commit 190c79dc39

Frequently asked questions

npx skillmds@latest add adu2021/pepo-token-level-multimodal-policy