Mebench Eval

Evaluates vision-language models' ability to ground objects and exhibit mutual exclusivity bias when mapping novel pseudo-labels to unknown items in cluttered scenes. It also measures spatial reasoning capabilities and the model's ability to resolve ambiguity among multiple novel objects. Use when the user wants to benchmark on MEBench, or asks about evaluating this task. Reports ME score.

qhjqhj00 2e458b9 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mebench-eval commit 2e458b9c0b

Frequently asked questions

npx skillmds add qhjqhj00/mebench-eval