Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

agentskillexchange Updated 28 repo stars

File contents

Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

Installation

Requirements and caveats from upstream:

  • N.B. Docker is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/apache-tika-content-extraction-hub commit 124dd1d80a

Frequently asked questions

npx skillmds@latest add agentskillexchange/apache-tika-content-extraction-hub