Mmdeepresearch Bench A Benchmark For Multimodal

Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where each task provides an image-text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emphas...

adu2021 Updated

File contents

Overview

This skill covers mmdeepresearch-bench: a benchmark for multimodal deep research agents. It addresses critical challenges in autonomous agent development.

Key Concepts

The paper introduces novel approaches to:

  • Agent evaluation and benchmarking
  • Improving agent efficiency and reasoning
  • Designing robust agent systems

When to Use

Use this when working on:

  • Agent-based systems and evaluation
  • Autonomous reasoning and planning
  • Multi-agent frameworks

When NOT to Use

  • Non-agent applications
  • Tasks requiring implementation code (see the paper)

References

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/mmdeepresearch-bench-a-benchmark-for-multimodal commit 1f85f0a5e1

Frequently asked questions

npx skillmds@latest add adu2021/mmdeepresearch-bench-a-benchmark-for-multimodal