Knapsack Rl Budget Allocation

Improve RL training for LLMs by dynamically allocating exploration budget (rollout count) to tasks based on their difficulty and current learning status. Solves the knapsack problem of maximizing gradient signal within fixed compute budget, increasing non-zero policy gradients by 20-40% and achieving 2-4 point performance gains.

adu2021 cba7f95 14.1 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/knapsack-rl-budget-allocation commit cba7f95505

Frequently asked questions

npx skillmds@latest add adu2021/knapsack-rl-budget-allocation