Tool List
1. pdf_to_markdown
Convert PDF documents to Markdown format, preserving document structure, formulas, tables, and images.
Description: Use MinerU to parse PDF documents and output in Markdown format, supporting OCR, formula recognition, table extraction, and other features.
Parameters:
file_path (string, required): Absolute path to the PDF file
output_dir (string, required): Absolute path to the output directory
backend (string, optional): Parsing backend, options: hybrid-auto-engine (default), pipeline, vlm-auto-engine
language (string, optional): OCR language code, such as en (English), ch (Chinese), ja (Japanese), etc., defaults to auto-detection
enable_formula (boolean, optional): Whether to enable formula recognition, defaults to true
enable_table (boolean, optional): Whether to enable table extraction, defaults to true
start_page (integer, optional): Start page number (starting from 0), defaults to 0
end_page (integer, optional): End page number (starting from 0), defaults to -1 meaning parse all pages
Return Value:
{
"success": true,
"output_path": "/path/to/output",
"markdown_content": "Converted Markdown content...",
"images": ["List of image paths"],
"tables": ["List of table information"],
"formula_count": 10
}
Examples:
python .claude/skills/pdf-process/script/pdf_parser.py \
'{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output"}}'
# Use specific backend
python .claude/skills/pdf-process/script/pdf_parser.py \
'{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "backend": "pipeline"}}'
# Parse specific pages
python .claude/skills/pdf-process/script/pdf_parser.py \
'{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "start_page": 0, "end_page": 5}}'
2. pdf_to_json
Convert PDF documents to JSON format, including detailed layout and structural information.
Description: Use MinerU to parse PDF documents and output in JSON format, containing structured information such as text blocks, images, tables, formulas, etc.
Parameters:
file_path (string, required): Absolute path to the PDF file
output_dir (string, required): Absolute path to the output directory
backend (string, optional): Parsing backend, options: hybrid-auto-engine (default), pipeline, vlm-auto-engine
language (string, optional): OCR language code, such as en (English), ch (Chinese), ja (Japanese), etc., defaults to auto-detection
enable_formula (boolean, optional): Whether to enable formula recognition, defaults to true
enable_table (boolean, optional): Whether to enable table extraction, defaults to true
start_page (integer, optional): Start page number (starting from 0), defaults to 0
end_page (integer, optional): End page number (starting from 0), defaults to -1 meaning parse all pages
Return Value:
{
"success": true,
"output_path": "/path/to/output.json",
"pages": [
{
"page_no": 0,
"page_size": [595, 842],
"blocks": [
{
"type": "text",
"text": "Text content",
"bbox": [x, y, x, y]
}
],
"images": [],
"tables": [],
"formulas": []
}
],
"metadata": {
"total_pages": 10,
"author": "Author",
"title": "Title"
}
}
Examples:
python .claude/skills/pdf-process/script/pdf_parser.py \
'{"name": "pdf_to_json", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output"}}'
# Use specific backend and language
python .claude/skills/pdf-process/script/pdf_parser.py \
'{"name": "pdf_to_json", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "backend": "hybrid-auto-engine", "language": "ch"}}'
Installation Instructions
1. Install MinerU
# Update pip and install uv
pip install --upgrade pip
pip install uv
# Install MinerU (including all features)
uv pip install -U "mineru[all]"
2. Verify Installation
# Check if MinerU is installed successfully
mineru --version
# Test basic functionality
mineru --help
3. System Requirements
- Python Version: 3.10-3.13
- Operating System: Linux / Windows / macOS 14.0+
- Memory:
- Using
pipeline backend: minimum 16GB, recommended 32GB+
- Using
hybrid/vlm backend: minimum 16GB, recommended 32GB+
- Disk Space: minimum 20GB (SSD recommended)
- GPU (optional):
pipeline backend: supports CPU-only
hybrid/vlm backend: requires NVIDIA GPU (Volta architecture and above) or Apple Silicon
Use Cases
- Academic Paper Parsing: Extract structured content such as formulas, tables, and images
- Technical Document Conversion: Convert PDF documents to Markdown for version control and online publishing
- OCR Processing: Process scanned PDFs and garbled PDFs
- Multilingual Documents: Supports OCR recognition for 109 languages
- Batch Processing: Batch convert multiple PDF documents
Backend Selection Recommendations
- hybrid-auto-engine (default): Balanced accuracy and speed, suitable for most scenarios
- pipeline: Suitable for CPU-only environments, best compatibility
- vlm-auto-engine: Highest accuracy, requires GPU acceleration
Notes
- File Paths: All paths must be absolute paths
- Output Directory: Non-existent directories will be created automatically
- Performance: Using GPU can significantly improve parsing speed
- Page Numbers: Page numbers start counting from 0
- Memory: Processing large documents may consume more memory
Troubleshooting
Common Issues
Installation Failure:
- Ensure using Python 3.10-3.13
- Windows only supports Python 3.10-3.12 (ray does not support 3.13)
- Using
uv pip install can resolve most dependency conflicts
Insufficient Memory:
- Use
pipeline backend
- Limit parsing pages:
start_page and end_page
- Reduce virtual memory allocation
Slow Parsing Speed:
- Enable GPU acceleration
- Use
hybrid-auto-engine backend
- Disable unnecessary features (formulas, tables)
Low OCR Accuracy:
- Specify the correct document language
- Ensure the backend supports OCR (use
pipeline or hybrid-*)
Related Resources
1---2name: pdf-process-mineru3description: PDF document parsing tool based on local MinerU, supports converting PDF to Markdown, JSON, and other machine-readable formats.4---56## Tool List78### 1. pdf_to_markdown910Convert PDF documents to Markdown format, preserving document structure, formulas, tables, and images.1112**Description**: Use MinerU to parse PDF documents and output in Markdown format, supporting OCR, formula recognition, table extraction, and other features.1314**Parameters**:15- `file_path` (string, required): Absolute path to the PDF file16- `output_dir` (string, required): Absolute path to the output directory17- `backend` (string, optional): Parsing backend, options: `hybrid-auto-engine` (default), `pipeline`, `vlm-auto-engine`18- `language` (string, optional): OCR language code, such as `en` (English), `ch` (Chinese), `ja` (Japanese), etc., defaults to auto-detection19- `enable_formula` (boolean, optional): Whether to enable formula recognition, defaults to true20- `enable_table` (boolean, optional): Whether to enable table extraction, defaults to true21- `start_page` (integer, optional): Start page number (starting from 0), defaults to 022- `end_page` (integer, optional): End page number (starting from 0), defaults to -1 meaning parse all pages2324**Return Value**:25```json26{27 "success": true,28 "output_path": "/path/to/output",29 "markdown_content": "Converted Markdown content...",30 "images": ["List of image paths"],31 "tables": ["List of table information"],32 "formula_count": 1033}34```3536**Examples**:37```bash38python .claude/skills/pdf-process/script/pdf_parser.py \39 '{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output"}}'4041# Use specific backend42python .claude/skills/pdf-process/script/pdf_parser.py \43 '{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "backend": "pipeline"}}'4445# Parse specific pages46python .claude/skills/pdf-process/script/pdf_parser.py \47 '{"name": "pdf_to_markdown", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "start_page": 0, "end_page": 5}}'48```4950---5152### 2. pdf_to_json5354Convert PDF documents to JSON format, including detailed layout and structural information.5556**Description**: Use MinerU to parse PDF documents and output in JSON format, containing structured information such as text blocks, images, tables, formulas, etc.5758**Parameters**:59- `file_path` (string, required): Absolute path to the PDF file60- `output_dir` (string, required): Absolute path to the output directory61- `backend` (string, optional): Parsing backend, options: `hybrid-auto-engine` (default), `pipeline`, `vlm-auto-engine`62- `language` (string, optional): OCR language code, such as `en` (English), `ch` (Chinese), `ja` (Japanese), etc., defaults to auto-detection63- `enable_formula` (boolean, optional): Whether to enable formula recognition, defaults to true64- `enable_table` (boolean, optional): Whether to enable table extraction, defaults to true65- `start_page` (integer, optional): Start page number (starting from 0), defaults to 066- `end_page` (integer, optional): End page number (starting from 0), defaults to -1 meaning parse all pages6768**Return Value**:69```json70{71 "success": true,72 "output_path": "/path/to/output.json",73 "pages": [74 {75 "page_no": 0,76 "page_size": [595, 842],77 "blocks": [78 {79 "type": "text",80 "text": "Text content",81 "bbox": [x, y, x, y]82 }83 ],84 "images": [],85 "tables": [],86 "formulas": []87 }88 ],89 "metadata": {90 "total_pages": 10,91 "author": "Author",92 "title": "Title"93 }94}95```9697**Examples**:98```bash99python .claude/skills/pdf-process/script/pdf_parser.py \100 '{"name": "pdf_to_json", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output"}}'101102# Use specific backend and language103python .claude/skills/pdf-process/script/pdf_parser.py \104 '{"name": "pdf_to_json", "arguments": {"file_path": "/path/to/document.pdf", "output_dir": "/path/to/output", "backend": "hybrid-auto-engine", "language": "ch"}}'105```106107---108109## Installation Instructions110111### 1. Install MinerU112113```bash114# Update pip and install uv115pip install --upgrade pip116pip install uv117118# Install MinerU (including all features)119uv pip install -U "mineru[all]"120```121122### 2. Verify Installation123124```bash125# Check if MinerU is installed successfully126mineru --version127128# Test basic functionality129mineru --help130```131132### 3. System Requirements133134- **Python Version**: 3.10-3.13135- **Operating System**: Linux / Windows / macOS 14.0+136- **Memory**:137 - Using `pipeline` backend: minimum 16GB, recommended 32GB+138 - Using `hybrid/vlm` backend: minimum 16GB, recommended 32GB+139- **Disk Space**: minimum 20GB (SSD recommended)140- **GPU** (optional):141 - `pipeline` backend: supports CPU-only142 - `hybrid/vlm` backend: requires NVIDIA GPU (Volta architecture and above) or Apple Silicon143144## Use Cases1451461. **Academic Paper Parsing**: Extract structured content such as formulas, tables, and images1472. **Technical Document Conversion**: Convert PDF documents to Markdown for version control and online publishing1483. **OCR Processing**: Process scanned PDFs and garbled PDFs1494. **Multilingual Documents**: Supports OCR recognition for 109 languages1505. **Batch Processing**: Batch convert multiple PDF documents151152## Backend Selection Recommendations153154- **hybrid-auto-engine** (default): Balanced accuracy and speed, suitable for most scenarios155- **pipeline**: Suitable for CPU-only environments, best compatibility156- **vlm-auto-engine**: Highest accuracy, requires GPU acceleration157158## Notes1591601. **File Paths**: All paths must be absolute paths1612. **Output Directory**: Non-existent directories will be created automatically1623. **Performance**: Using GPU can significantly improve parsing speed1634. **Page Numbers**: Page numbers start counting from 01645. **Memory**: Processing large documents may consume more memory165166## Troubleshooting167168### Common Issues1691701. **Installation Failure**:171 - Ensure using Python 3.10-3.13172 - Windows only supports Python 3.10-3.12 (ray does not support 3.13)173 - Using `uv pip install` can resolve most dependency conflicts1741752. **Insufficient Memory**:176 - Use `pipeline` backend177 - Limit parsing pages: `start_page` and `end_page`178 - Reduce virtual memory allocation1791803. **Slow Parsing Speed**:181 - Enable GPU acceleration182 - Use `hybrid-auto-engine` backend183 - Disable unnecessary features (formulas, tables)1841854. **Low OCR Accuracy**:186 - Specify the correct document language187 - Ensure the backend supports OCR (use `pipeline` or `hybrid-*`)188189## Related Resources190191- MinerU Official Documentation: https://opendatalab.github.io/MinerU/192- MinerU GitHub: https://github.com/opendatalab/MinerU193- Online Demo: https://mineru.net/