# MinerU PDF-to-Markdown Document Parser

> Transforms complex PDFs into LLM-ready markdown and JSON using MinerU, a high-accuracy document intelligence pipeline. Extracts text, tables, formulas, and images from scientific papers, reports, and scanned documents with layout-aware parsing.

- Skill: `agentskillexchange/mineru-pdf-to-markdown-document-parser` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/mineru-pdf-to-markdown-document-parser`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/mineru-pdf-to-markdown-document-parser/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/mineru-pdf-to-markdown-document-parser

---


# MinerU PDF-to-Markdown Document Parser

Transforms complex PDFs into LLM-ready markdown and JSON using MinerU, a high-accuracy document intelligence pipeline. Extracts text, tables, formulas, and images from scientific papers, reports, and scanned documents with layout-aware parsing.

## Installation

Use the upstream install or setup path that matches your environment:
- pip install --upgrade pip
- pip install uv
- uv pip install -U "mineru[all]"
- git clone https://github.com/opendatalab/MinerU.git

Requirements and caveats from upstream:
- [![PyPI - Python Version](https://img.shields.io/pypi/pyversions/mineru)](https://pypi.org/project/mineru/)
- | Development | Python / Go / TypeScript SDK · CLI · REST API · Docker |
- The official online version has the same functionality as the client, with a beautiful interface and rich features, requires login to use

Basic usage or getting-started notes:
- While maintaining high accuracy, it keeps resource usage extremely low and continues to support inference in pure CPU environments.
- Optimized the parsing pipeline with a sliding-window mechanism, significantly reducing peak memory usage in long-document scenarios, so documents with tens of thousands of pages no longer need to be split manually.
- This update is not just a set of feature enhancements, but a key leap forward in MinerU's overall system capabilities. We specifically addressed the peak memory usage issue in long-document parsing. Through optimizati...

- Source: https://github.com/opendatalab/MinerU
- Extracted from upstream docs: https://raw.githubusercontent.com/opendatalab/MinerU/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/mineru-pdf-to-markdown-document-parser/)

