# Apache Tika Content Extraction Hub

> Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

- Skill: `agentskillexchange/apache-tika-content-extraction-hub` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/apache-tika-content-extraction-hub`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/apache-tika-content-extraction-hub/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/apache-tika-content-extraction-hub

---


# Apache Tika Content Extraction Hub

Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.

## Installation

Requirements and caveats from upstream:
- **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:
- ===========
- **Parse a file in Java:**
- java

- Source: https://github.com/apache/tika
- Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/)

