Innovator Vl Eval

Evaluates multimodal large language models across general vision, mathematical reasoning, and specialized scientific domains to measure visual perception, instruction following, and domain-specific knowledge retention. Use when the user wants to benchmark on AI2D, OCRBench, ChartQA, MMMU(Val), MMMU-Pro (Standard), MMStar, VStar-Bench, MMBench-EN, MME-RealWorld, DocVQA(Val), InfoVQA(Val), SEED-Bench, SEED-Bench-2-plus, RealWorldQA, MathVision, MathVerse, MathVista, WeMath, ScienceQA, RxnBench, MolParse, OpenRxn, EMVista, SuperChem, SmolInstruct, ProteinLMBench, SFE, MicroVQA, MSEarth-MCQ, XLRS-Bench-lite, or asks about evaluating this task. Reports accuracy.

qhjqhj00 335a487 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/innovator-vl-eval commit 335a487232

Frequently asked questions

npx skillmds add qhjqhj00/innovator-vl-eval