Deploy document-to-JSON extraction APIs and ETL pipelines with Unstract
Use Unstract when an operator needs agents or automation pipelines to turn recurring PDFs, scans, and document batches into structured JSON through prompt-defined extraction workflows.
Prerequisites
Unstract platform; Docker and Docker Compose; Git; LLM provider credentials; target document source; optional API, ETL, MCP, or n8n integration
Installation
Use the upstream install or setup path that matches your environment:
- git clone https://github.com/Zipstack/unstract.git
Requirements and caveats from upstream:
- n8n Node — Drop into existing automation workflows. Docs →
Basic usage or getting-started notes:
| Deployment | Custom infrastructure | ./run-platform.sh or managed cloud |
./run-platform.sh
./run-platform.sh -v v0.1.0
Extracted from upstream docs: https://raw.githubusercontent.com/Zipstack/unstract/HEAD/README.md