Prl Process Reward Learning

Improves LLM reasoning by decomposing RL objectives into intermediate process rewards assigned to reasoning steps, improving both final accuracy and reasoning capacity without expensive Monte Carlo Tree Search.

adu2021 edba4b0 9.6 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/prl-process-reward-learning commit edba4b07b2

Frequently asked questions

npx skillmds@latest add adu2021/prl-process-reward-learning