Mint 1t Eval

Evaluates the multimodal interleaved reasoning and in-context learning capabilities of large multimodal models (LMMs) across image captioning, visual question answering, and multi-image reasoning tasks. Use when the user wants to benchmark on COCO (Karpathy test), TextCaps, VQAv2, OK-VQA, TextVQA, VizWiz, MMMU, Mantis-Eval, or asks about evaluating this task. Reports scores.

qhjqhj00 d09a573 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mint-1t-eval commit d09a573ac7

Frequently asked questions

npx skillmds add qhjqhj00/mint-1t-eval