Agent Reasoning Reward Model

Build multi-faceted reward models for agent trajectories that provide structured feedback on intermediate reasoning quality. Implement explicit reasoning traces, focused critiques with refinement guidance, and overall process scores to train more effective agentic agents without relying solely on sparse outcome rewards.

adu2021 4a54152 5.8 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/agent-reasoning-reward-model commit 4a54152062

Frequently asked questions

npx skillmds@latest add adu2021/agent-reasoning-reward-model