Prepare agent-ready PDF and document extraction with PyMuPDF
Use PyMuPDF to extract text, layout metadata, tables, images, Markdown, or JSON from PDFs and documents before an agent ingests or reviews them.
Prerequisites
Python 3.10+, PyMuPDF, source PDFs or supported document files, and optional pymupdf4llm or Tesseract OCR for LLM-ready Markdown or scanned pages.
Installation
Use the upstream install or setup path that matches your environment:
- pip install pymupdf
- pip install pymupdf-fonts
- pip install pymupdf4llm
- pip install pymupdfpro
Requirements and caveats from upstream:
- [](https://demo.pymupdf.io?utm_source=github&utm_medium=referral&utm_campaign=pymupdf_github&utm_content=badges&utm_t...
Basic usage or getting-started notes:
| Package | Purpose |
|---|---|
| pymupdf-fonts | Extended font collection for text output |
Extracted from upstream docs: https://raw.githubusercontent.com/pymupdf/PyMuPDF/HEAD/README.md