Toolbench Eval

This benchmark evaluates an LLM's ability to understand API documentation, select the correct tool, and generate valid arguments to fulfill a user's goal. It probes tool manipulation capabilities across a diverse set of 8 real-world applications, ranging from simple single-call tasks to complex multi-step reasoning. Use when the user wants to benchmark on ToolBench, or asks about evaluating this task. Reports success rate.

qhjqhj00 21e62cd 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/toolbench-eval commit 21e62cd38e

Frequently asked questions

npx skillmds add qhjqhj00/toolbench-eval