Noisytoolbench Eval

This benchmark evaluates how well LLM agents handle ambiguous or unclear user instructions by measuring their ability to ask clarifying questions, execute correct tool calls, and generate accurate final answers. It also assesses interaction efficiency by tracking redundant questions and total action steps. Use when the user wants to benchmark on NoisyToolBench, or asks about evaluating this task. Reports A1.

qhjqhj00 a3e9528 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/noisytoolbench-eval commit a3e95282a5

Frequently asked questions

npx skillmds add qhjqhj00/noisytoolbench-eval