1183 Agentbench E0f23d17

AgentBench benchmark for evaluating LLMs as autonomous agents across diverse environments including OS, databases, and web.

tools-only Updated 7 repo stars

File contents

tools-only/X-Skills/tree/main/commercial/1183-agentbench_e0f23d17 commit 5270b94913

Frequently asked questions

npx skillmds@latest add tools-only/1183-agentbench-e0f23d17