Multimodal Tool Use Eval

Evaluates the ability of agentic multimodal models to perform visual perception, document understanding, and mathematical reasoning. It specifically probes whether models can strategically decide when to invoke external tools (e.g., image cropping, web search, Python code execution) versus answering directly, balancing task accuracy with tool efficiency. Use when the user wants to benchmark on V-Bench, HRBench-4K/8K, TreeBench, MME-RealWorld, SEEDBench2-Plus, CharXiv, MathVista_mini, MathVerse_mini, WeMath, DynaMath, LogicVista, or asks about evaluating this task. Reports accuracy.

qhjqhj00 563279d 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-tool-use-eval commit 563279d191

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-tool-use-eval