Parse documents into structured content for agent ingestion with Dedoc
Extract document text, tables, logical structure, and metadata into normalized output before RAG, review, or workflow automation.
Prerequisites
Dedoc Docker service or Python library, source documents
Installation
Use the upstream install or setup path that matches your environment:
- docker pull dedocproject/dedoc
- docker run -p 1231:1231 --rm dedocproject/dedoc python3 /dedoc_root/dedoc/main.py
- git clone https://github.com/ispras/dedoc
- docker compose up --build
Requirements and caveats from upstream:
- Dedoc is implemented in Python and works with semi-structured data formats (DOC/DOCX, ODT, XLS/XLSX, CSV, TXT, JSON) and unstructured data formats like images (PNG, JPG etc.), archives (ZIP, RAR etc.), PDF and HTML fo...
Basic usage or getting-started notes:
Source: https://github.com/ispras/dedoc
Extracted from upstream docs: https://raw.githubusercontent.com/ispras/dedoc/HEAD/README.md