Unstructured Document Partitioning and ETL Library for LLM Pipelines
Unstructured is an open-source library for ingesting and partitioning PDFs, HTML, Office documents, emails, and other unstructured inputs into structured elements and metadata. It is commonly used as a preprocessing layer for RAG, search, extraction, and downstream AI pipelines.
Prerequisites
Python 3.11+
Installation
Use the upstream install or setup path that matches your environment:
- docker pull downloads.unstructured.io/unstructured-io/unstructured:latest
- docker run -dt --name unstructured downloads.unstructured.io/unstructured-io/unstructured:latest
- docker exec -it unstructured bash
- make docker-build
Requirements and caveats from upstream:
Basic usage or getting-started notes:
:eight_pointed_black_star: Quick Start
Run the library in a container
Extracted from upstream docs: https://raw.githubusercontent.com/Unstructured-IO/unstructured/HEAD/README.md