AI agent skills for PDF and document handling
Document skills cover the work that looks trivial and is not: extracting text from a PDF that has columns, filling a form, producing a .docx that opens correctly in Word. The best of them are explicit about the library they drive and what they cannot do — scanned pages need OCR, and a skill that pretends otherwise will quietly return nothing. If the output is going to a person rather than a parser, check whether the skill handles layout or only text.
People land here searching for “claude code pdf skill”, “agent skill to extract text from pdf”, “docx generation skill”.
117 PDF and document handling skills Refine in search →
-
majiayu000 Bundle Document ConverterConvert between Markdown, DOCX, and PDF formats bidirectionally. Handles text extraction from PDF/DOCX, markdown to document conversion. Use when converting document formats or extracting structured content from Word or PDF files.
567 -
neuromechanist Bundle Document ProcessingThis skill should be used when the user says "process documents", "extract text from PDF", "OCR this document", "convert PDF to markdown", "extract emails from documents", "parse document", "document conversion", "batch OCR", "extract structured data from PDF", "read PDF", "extract tables from PDF", "convert Word document", "convert docx to markdown", or wants to extract, convert, or process documents and scanned images.
-
diegosouzapw Bundle DocumentsRead, write, convert, and analyze documents — routes to PDF, DOCX, XLSX, PPTX sub-skills for creation, editing, extraction, and format conversion. USE WHEN document, process file, create document, convert format, extract text, PDF, DOCX, XLSX, PPTX, Word, Excel, spreadsheet, PowerPoint, presentation, slides, consulting report, large PDF, merge PDF, fill form, tracked changes, redlining.
54 -
akaihola Bundle Read As MarkdownConvert binary document files (PDF, DOCX, EPUB) to cached Markdown for reading. Use when the user asks to "read a PDF", "convert docx to markdown", "extract text from a document", "show me this PDF as markdown", "read this paper", or mentions reading, viewing, or extracting text from .pdf, .docx, or .epub files. This skill is for READING documents as markdown — not for creating or editing them (use the pdf/docx skills for that).
-
majiayu000 Bundle Document Processor 2Extract and process content from PDFs and DOCX files. Handles large files, OCR for scanned documents, page splitting, and markdown conversion. Use when: (1) Processing PDF references in notes, (2) Extracting text from large documents for analysis, (3) Converting DOCX to markdown, (4) Handling scanned/image PDFs with OCR, (5) Integrating with Obsidian or note-taking workflows, (6) Splitting large documents into manageable chunks. Invoke with: /process-document, /extract-pdf, /extract-docx, or say "use document-processor skill to..."
567 -
oyi77 Bundle Book To SkillConverts technical books and documents (PDF, EPUB, DOCX, HTML, Markdown, RTF, MOBI) into structured agent skills with frameworks, mental models, chapter references, and decision rules. Includes a full extraction pipeline for turning owned documents into reusable skills.
10 -
seaworld008 Bundle Markdown ToolsConvert PDF, DOCX, PPTX, and other documents to Markdown, preserving tables, images, and structure with the appropriate extraction tool.
65 -
tbusos Bundle Doc To MarkdownConvert documents (PDF, DOCX) in a directory to well-formatted Markdown files with images extracted. Use this skill whenever the user wants to convert documents to markdown, make documents readable for future reference, extract content from PDFs or Word files into markdown format, or batch-convert a folder of documents. Trigger on phrases like "convert to markdown", "转成markdown", "转换成md", "文档转换", "把文档变成可读的", "提取文档内容", or any request involving turning PDF/DOCX files into a text-based readable format with images preserved. Not for the reverse direction: to turn Markdown into PDF, use md-to-pdf instead.
-
dvcrn Bundle ZeroxConvert documents (PDF, DOCX, PPTX, images, etc.) to Markdown using the zerox library. Use when the user needs to extract text content from document files.
32 -
modbender Bundle ZeroxConvert documents (PDF, DOCX, PPTX, images, etc.) to Markdown using the zerox library. Use when the user needs to extract text content from document files.
12 -
diegosouzapw Bundle DOCX To PDFConvert Word documents (DOCX files) to PDF format in batch. Use when the user uploads one or more Word documents and asks to convert them to PDF, export to PDF, or save as PDF. The tool processes multiple files efficiently and provides a summary of successful conversions and any failures.
54 -
majiayu000 Bundle Document Conversion将 DOC/DOCX/PDF/PPT/PPTX 文档转换为 Markdown 格式。自动检测 PDF 类型(电子版/扫描版),提取图片到独立目录。当管理员入库非 Markdown 文档时使用此 Skill。触发条件:入库 DOC/DOCX/PDF/PPT/PPTX 格式文件。
567 -
gabrielmoreira Skill Ocr And Documents 2Extract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
17 -
modbender Bundle Office To Markdown Converter Skill V2Convert office documents (PDF, DOC, DOCX, PPTX) to Markdown format. This skill uses the word-extractor library for .doc support and provides full OpenClaw integration.
12 -
chunpu Bundle Doc To TxtConvert DOC, DOCX, and PDF files to TXT format. Invoke when user wants to extract text from these document types.
-
comeonoliver Skill Markdown ExporterMarkdown Exporter
61 -
braxtonrose4 Bundle Ocr And DocumentsExtract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
-
intertwine Bundle Ocr And DocumentsExtract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
-
vamseeachanta Bundle Ocr And DocumentsExtract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
-
vamseeachanta Bundle Ocr And Documents 2Extract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
-
robomotionio Bundle Ocr And DocumentsExtract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
-
thechandanbhagat Skill DOCXWork with Word documents (DOCX/DOC) - read, extract content, create documents, convert formats, and process templates. Use when the user asks to work with Word files or .docx documents.
-
cas-bigdatalab Bundle DOCX Text Extract从DOCX文件中提取文本内容
-
johnalbertini14-glitch Bundle ZeroxConvert documents (PDF, DOCX, PPTX, images, etc.) to Markdown using the zerox library. Use when the user needs to extract text content from document files.
1
How to install a PDF and document handling skill
- Compare the skills. Read the PDF and document handling skills below — each page shows the full SKILL.md, its safety verdict, and the capabilities it declares.
- Install it. Run npx skillmds@latest add <owner>/<name>. The CLI writes the skill into every agent directory it detects, or use --agent to pin one.
- Use it. Restart your agent. It loads the skill on demand the next time you ask for something that matches — you do not have to name the skill.
What makes a good PDF and document handling skill?
Document skills cover the work that looks trivial and is not: extracting text from a PDF that has columns, filling a form, producing a .docx that opens correctly in Word. The best of them are explicit about the library they drive and what they cannot do — scanned pages need OCR, and a skill that pretends otherwise will quietly return nothing. If the output is going to a person rather than a parser, check whether the skill handles layout or only text.
Every skill listed here is a plain SKILL.md file in the format Anthropic documents for Agent Skills, read unchanged by Cursor, OpenAI Codex and 60+ agents. Each passes a safety review before it is publicly listed; the verdict and declared capabilities are on every skill's page.
Frequently asked questions
What is an agent skill for PDF and document handling?
Document skills cover the work that looks trivial and is not: extracting text from a PDF that has columns, filling a form, producing a .docx that opens correctly in Word. The best of them are explicit about the library they drive and what they cannot do — scanned pages need OCR, and a skill that pretends otherwise will quietly return nothing. If the output is going to a person rather than a parser, check whether the skill handles layout or only text.
Do document skills work on scanned PDFs?
Only if they invoke OCR, and most do not by default. A scanned page is an image; a text-extraction skill will find no text and should say so rather than guessing. Skills with OCR support name the engine they use.
Which agents can use these PDF and document handling skills?
Any agent that reads SKILL.md files: Claude Code, Claude.ai, Cursor, OpenAI Codex, Windsurf, OpenCode and 60+ others. The format is not vendor-specific, so the same file works everywhere — each agent just keeps its skills in a different directory, listed at /agents.
Are these PDF and document handling skills free?
Yes. Searching, reading and installing skills on SkillMD is free and needs no account. Individual skills carry their own licence, shown on each skill's page.