Convert PDFs and document images into agent-ready Markdown with Docext
Run Docext locally or with a chosen model backend to turn PDFs and document images into structured Markdown for RAG, extraction, and review workflows.
Prerequisites
Python 3.11 environment, Docext package, source PDFs or document images, and a supported VLM backend such as vLLM, Ollama, or a configured hosted model provider
Installation
Basic usage or getting-started notes:
On-premises deployment: Run entirely on your own infrastructure (Linux, MacOS)
For more details (Installation, Usage, and so on), please check out the feature guide.
Extracted from upstream docs: https://raw.githubusercontent.com/NanoNets/docext/HEAD/README.md