Toolrm Outcome Reward Tool Calling

Specialized outcome reward models (1.7B-14B) for evaluating tool-calling performance in LLMs. Addresses the critical gap in reward modeling where general-purpose models miss key signals of effective tool use. Enables better Best-of-N sampling, data filtering, and RL-based policy training through FC-RewardBench evaluation framework.

adu2021 be496d6 21.9 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/toolrm-outcome-reward-tool-calling commit be496d65c3

Frequently asked questions

npx skillmds@latest add adu2021/toolrm-outcome-reward-tool-calling