Gsm8k V Eval

This benchmark evaluates vision-language models' ability to perform multi-step mathematical reasoning using purely visual, comic-style narratives instead of text. It specifically probes challenges in inter-image semantic understanding, object grounding, and extracting numerical relationships from multi-panel visual contexts. Use when the user wants to benchmark on GSM8K-V, or asks about evaluating this task. Reports accuracy.

qhjqhj00 eeff98f 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gsm8k-v-eval commit eeff98f5eb

Frequently asked questions

npx skillmds add qhjqhj00/gsm8k-v-eval