Agentrewardbench Eval

This benchmark evaluates the effectiveness of LLM-based judges in automatically assessing web agent trajectories. It probes the judges' ability to correctly predict task success, detect side effects, and identify repetitive actions by comparing their outputs against expert human annotations. Use when the user wants to benchmark on AgentRewardBench, or asks about evaluating this task. Reports precision.

qhjqhj00 c75b337 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agentrewardbench-eval commit c75b337136

Frequently asked questions

npx skillmds add qhjqhj00/agentrewardbench-eval