Mmdr Bench Eval

This benchmark evaluates a model's ability to handle complex, multi-turn visually-grounded dialogue and follow intricate instructions. It probes sustained contextual understanding, visual entity tracking across turns, and multi-step reasoning depth in dynamic multi-modal interactions. Use when the user wants to benchmark on MMDR-Bench, or asks about evaluating this task. Reports average human evaluation ratings.

qhjqhj00 4d410c5 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmdr-bench-eval commit 4d410c5922

Frequently asked questions

npx skillmds add qhjqhj00/mmdr-bench-eval