Mm Vet V2 Eval

This benchmark evaluates large multimodal models on integrated vision-language capabilities, with a specific focus on sequential image-text understanding, spatial reasoning, knowledge retrieval, and long-form generation. It probes how well models can process interleaved visual and textual inputs to answer complex, real-world questions. Use when the user wants to benchmark on MM-Vet v2, or asks about evaluating this task. Reports MM-Vet-v2 score.

qhjqhj00 ea2ec8a 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-vet-v2-eval commit ea2ec8ac4f

Frequently asked questions

npx skillmds add qhjqhj00/mm-vet-v2-eval