Arxivcap Eval

Evaluates large vision-language models' ability to comprehend and generate text for scientific figures. It probes capabilities in single and multi-figure captioning, contextualized captioning using in-context examples, and inferring paper titles from figure-caption sequences. Use when the user wants to benchmark on ArXivCap, or asks about evaluating this task. Reports BLEU-2.

qhjqhj00 2a038f1 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/arxivcap-eval commit 2a038f133f

Frequently asked questions

npx skillmds add qhjqhj00/arxivcap-eval