Vlmbench Eval

This benchmark evaluates a robot agent's ability to execute 6D manipulation tasks guided by natural language instructions and visual observations. It probes compositional reasoning, object localization, and precise pose estimation in both seen and unseen object settings. Use when the user wants to benchmark on VLMbench, or asks about evaluating this task. Reports success rate.

qhjqhj00 3ccc58d 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vlmbench-eval commit 3ccc58dda4

Frequently asked questions

npx skillmds add qhjqhj00/vlmbench-eval