Apache Tika Document Extractor

Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.

agentskillexchange Updated 28 repo stars

File contents

Apache Tika Document Extractor

Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.

Installation

Requirements and caveats from upstream:

  • N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/apache-tika-document-extractor commit 5cf4bb6e3b

Frequently asked questions

npx skillmds@latest add agentskillexchange/apache-tika-document-extractor