Mir Eval

Evaluates multimodal large language models on progressive, interleaved multi-image reasoning tasks. It probes the model's ability to perform structured, step-by-step reasoning across multiple images, including text-to-region alignment, cross-image relationship modeling, and analytical inference. Use when the user wants to benchmark on MIR, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6e9319b 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mir-eval commit 6e9319b00b

Frequently asked questions

npx skillmds add qhjqhj00/mir-eval