Tika Analyst

Expert skill for document intelligence tasks powered by a self-hosted Apache Tika instance. Use this skill whenever the agent needs to: extract text from any file (PDF, DOCX, PPTX, XLSX, email, EPUB, archive, image, video, audio, geospatial, or any other format); detect MIME types or file metadata; recursively unpack compound documents like ZIP archives, EML emails with attachments, or embedded Office OLE objects; run OCR on scanned PDFs or images; detect document language; audit documents for embedded macros; or extract raw embedded assets such as images and charts from Word docs or PDFs. Also trigger for any workflow that needs to triage unknown file types before deciding how to process them, or that needs to feed clean text into a downstream LLM. Always prefer this skill over ad-hoc file parsing — Tika handles 1,400+ MIME types deterministically and Tesseract OCR is included in the deployed full container.

mister2d Updated

File contents

mister2d/agent-skills/tree/main/tika-analyst commit a30bab052e

Frequently asked questions

npx skillmds@latest add mister2d/tika-analyst