Apache Tika Document Parser

Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.

agentskillexchange Updated 28 repo stars

File contents

Apache Tika Document Parser

Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.

Installation

Requirements and caveats from upstream:

  • N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/apache-tika-document-parser commit 7137052b3d

Frequently asked questions

npx skillmds@latest add agentskillexchange/apache-tika-document-parser