Vtc Bench Eval

Evaluates multimodal large language models' ability to perform agentic visual reasoning by composing multiple OpenCV-based tool calls. It probes long-horizon planning, precise tool selection, and the capacity to chain coarse- to fine-grained visual operations to solve complex multi-step problems. Use when the user wants to benchmark on VTC-Bench, or asks about evaluating this task. Reports Average Pass Rate (APR).

qhjqhj00 3dab823 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vtc-bench-eval commit 3dab82393f

Frequently asked questions

npx skillmds add qhjqhj00/vtc-bench-eval