Milebench Eval

Evaluates Multimodal Large Language Models (MLLMs) on long-context, multi-image comprehension. It probes capabilities like needle-in-a-haystack retrieval, image retrieval, temporal reasoning across multiple images, and semantic understanding in long multimodal contexts. Use when the user wants to benchmark on MileBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 0330abc 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/milebench-eval commit 0330abcca4

Frequently asked questions

npx skillmds add qhjqhj00/milebench-eval