Mm Bright Eval

This benchmark evaluates reasoning-intensive retrieval capabilities across text-only and multimodal settings. It probes models' ability to align visual and textual information, navigate technical domain queries, and rank relevant documents or images based on complex, multi-modal prompts. Use when the user wants to benchmark on MM-BRIGHT, or asks about evaluating this task. Reports nDCG@10.

qhjqhj00 58979e5 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-bright-eval commit 58979e5db6

Frequently asked questions

npx skillmds add qhjqhj00/mm-bright-eval