Exploring Reasoning Reward Agents

Apply Agent Reasoning Reward Model (Agent-RRM) structured critique to improve multi-step agent trajectories. Evaluates tool-use chains with explicit reasoning traces, focused critiques, and process scores. Use this skill when: - "Critique this agent's reasoning trace" - "Evaluate my tool-calling workflow and find flaws" - "Score this multi-step agent trajectory" - "Help me build a reward model for agent training" - "Improve this agent's reasoning with structured feedback" - "Debug why my agent pipeline produces wrong answers"

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/exploring-reasoning-reward-agents commit 25a6b114c2

Frequently asked questions

npx skillmds@latest add ndpvt-web/exploring-reasoning-reward-agents