Spiqa Eval

Evaluates multimodal long-context reasoning and figure/table comprehension on scientific papers. Tests direct question answering with images, full paper context, and chain-of-thought retrieval capabilities. Use when the user wants to benchmark on SPIQA, or asks about evaluating this task. Reports L3Score.

qhjqhj00 a308eff 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/spiqa-eval commit a308eff3f4

Frequently asked questions

npx skillmds add qhjqhj00/spiqa-eval