Agentbench Eval

Evaluates LLMs as autonomous agents across eight diverse, real-world environments requiring multi-turn interaction, long-term reasoning, decision-making, and strict instruction following. The benchmark measures success rates across code, game, and web-based tasks to identify performance gaps between commercial and open-source models. Use when the user wants to benchmark on AgentBench, or asks about evaluating this task. Reports overall_score.

qhjqhj00 dd76a0e 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agentbench-eval commit dd76a0ea64

Frequently asked questions

npx skillmds add qhjqhj00/agentbench-eval