Toolsandbox Eval

Evaluates LLM tool-use capabilities in a stateful, conversational, and interactive setting. It probes the model's ability to handle implicit state dependencies, canonicalize arguments, handle insufficient information, and maintain efficiency across single/multiple tool calls and user turns. Use when the user wants to benchmark on ToolSandbox, or asks about evaluating this task. Reports average similarity score.

qhjqhj00 6d06ae0 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/toolsandbox-eval commit 6d06ae0b10

Frequently asked questions

npx skillmds add qhjqhj00/toolsandbox-eval