Extract OCR-ready Markdown from documents with Zerox
Use Zerox to convert PDFs, images, and office documents into Markdown or structured extraction outputs using vision models.
Prerequisites
Node.js or Python Zerox package, graphicsmagick/ghostscript for Node PDF conversion or poppler for Python, model provider credentials
Installation
Use the upstream install or setup path that matches your environment:
- npm install zerox
- pip install py-zerox
Requirements and caveats from upstream:
- Zerox is available as both a Node and Python package.
- Node README - npm package
- Python README - pip package
Basic usage or getting-started notes:
| ------------------------- | ---------------------------- | -------------------------- |
| Image Processing | ✓ | ✓ |
| OpenAI Support | ✓ | ✓ |
Extracted from upstream docs: https://raw.githubusercontent.com/getomni-ai/zerox/HEAD/README.md