Scivqa Eval

Evaluates multimodal LLMs on closed-ended visual and non-visual question answering over scientific figures. It probes recognition of visual attributes (color, shape, position) and reasoning capabilities across diverse chart types. Use when the user wants to benchmark on SciVQA, or asks about evaluating this task. Reports ROUGE-1 F1.

qhjqhj00 7934a4c 2.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scivqa-eval commit 7934a4cf2f

Frequently asked questions

npx skillmds add qhjqhj00/scivqa-eval