Deep Search Mcts Rlvr Training

Overcome exploration bottlenecks in reasoning RL by integrating Monte Carlo Tree Search during training (not just inference). Global frontier selection and entropy-guided sampling reduce GPU hours by 5.7x while improving performance.

adu2021 Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/deep-search-mcts-rlvr-training commit 05a26646e0

Frequently asked questions

npx skillmds@latest add adu2021/deep-search-mcts-rlvr-training