Demystifying Rl Tool Agents

Comprehensive recipe for RL-training tool-using agents spanning reward design, data synthesis, model scaling, and algorithm selection. Seven ranked findings: scale-dependent rewards (curriculum for 1.5B–3B; dense for 7B), semi-sparse 'Macro' rewards balance specialization/transfer, 1K-sample sweet spot with 4:3:3 difficulty mix. Achieves SOTA on TravelPlanner with smaller models than leading proprietary systems.

adu2021 0562968 8.6 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.3-claude-opus-4.6/demystifying-rl-tool-agents commit 0562968327

Frequently asked questions

npx skillmds@latest add adu2021/demystifying-rl-tool-agents