# Apache Tika Document Parser

> Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.

- Skill: `agentskillexchange/apache-tika-document-parser` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/apache-tika-document-parser`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/apache-tika-document-parser/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/apache-tika-document-parser

---


# Apache Tika Document Parser

Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.

## Installation

Requirements and caveats from upstream:
- **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:
- ===========
- **Parse a file in Java:**
- java

- Source: https://github.com/apache/tika
- Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-parser/)

