Vlm Subtlebench Eval

This benchmark evaluates vision-language models' ability to perform subtle comparative reasoning between pairs of images. It probes capabilities across ten fine-grained difference types, including spatial, temporal, viewpoint, attribute, and existence changes, requiring models to detect and explain nuanced visual discrepancies that are often missed by standard prompting or simple image concatenation. Use when the user wants to benchmark on VLM-SubtleBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 862d176 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vlm-subtlebench-eval commit 862d1768b6

Frequently asked questions

npx skillmds add qhjqhj00/vlm-subtlebench-eval