Onethinker Eval

Evaluates a unified multimodal reasoning model's ability to perform visual understanding tasks across both static images and videos. It probes capabilities in question answering, captioning, spatial and temporal grounding, object tracking, and segmentation. Use when the user wants to benchmark on MMMU, MathVista, MathVerse, MMBench, MMStar, ScienceQA, AI2D, MMT-Bench, VideoMMMU, MMVU, VideoMME, VideoHolmes, LongVideoBench, LongVideo-Reason, VideoMathQA, MMSci-Caption, MMT-Caption, VideoMMLU-Caption, Charades, ActivityNet, ANet-RTL, RefCOCO, RefCOCO+, RefCOCOg, STVG, GOT-10k, MeViS, ReasonVOS, or asks about evaluating this task. Reports accuracy.

qhjqhj00 74d0ba8 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/onethinker-eval commit 74d0ba861d

Frequently asked questions

npx skillmds add qhjqhj00/onethinker-eval