# Mmdeepresearch Bench A Benchmark For Multimodal

> Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where each task provides an image-text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emphas...

- Skill: `adu2021/mmdeepresearch-bench-a-benchmark-for-multimodal` (Agent Skill)
- Install (CLI): `npx skillmds@latest add adu2021/mmdeepresearch-bench-a-benchmark-for-multimodal`
- Raw SKILL.md: https://api.skillmd.com/api/skills/adu2021/mmdeepresearch-bench-a-benchmark-for-multimodal/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- License: MIT
- Author: adu2021 (https://skillmd.com/u/adu2021)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/adu2021/mmdeepresearch-bench-a-benchmark-for-multimodal

---


## Overview

This skill covers mmdeepresearch-bench: a benchmark for multimodal deep research agents. It addresses critical challenges in autonomous agent development.

## Key Concepts

The paper introduces novel approaches to:
- Agent evaluation and benchmarking
- Improving agent efficiency and reasoning
- Designing robust agent systems

## When to Use

Use this when working on:
- Agent-based systems and evaluation
- Autonomous reasoning and planning
- Multi-agent frameworks

## When NOT to Use

- Non-agent applications
- Tasks requiring implementation code (see the paper)

## References

- Paper: https://arxiv.org/abs/2601.12346
- PDF: https://arxiv.org/pdf/2601.12346

