Nemotron Nano V2 Vl Eval

Evaluates a 12B vision-language model's capabilities across multimodal understanding, long-context reasoning, document/OCR processing, video comprehension, and pure text reasoning. It probes the model's ability to handle diverse visual inputs, follow instructions, and perform complex STEM and code reasoning under varying decoding and reasoning budget constraints. Use when the user wants to benchmark on MMBench V1.1, MMMU, OCRBench, DocVQA, LongVideoBench, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.

qhjqhj00 0e9e92c 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/nemotron-nano-v2-vl-eval commit 0e9e92c3a7

Frequently asked questions

npx skillmds add qhjqhj00/nemotron-nano-v2-vl-eval