Tool Learning Eval

Evaluates foundation models' ability to decompose complex instructions, reason over subgoals, and dynamically select/call external APIs or tools to complete tasks across diverse domains like translation, mathematics, web search, and data processing. Use when the user wants to benchmark on MLQA, ASDiv, MathQA, RealTimeQA, HotpotQA, WebShop, ALFWorld, Curated (Map), Curated (Weather), Curated (Stock), Curated (Slides), Curated (Tables), Curated (KGs), Curated (Cooking), Curated (Movie), Curated (AI Painting), Curated (3D Model Construction), Curated (Chemical Properties), Curated (Database), or asks about evaluating this task. Reports accuracy / success rate.

qhjqhj00 a44c515 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tool-learning-eval commit a44c515864

Frequently asked questions

npx skillmds add qhjqhj00/tool-learning-eval