Pdf2txt

This skill should be used when users need to extract text content from PDF files. It provides robust text extraction from PDF documents using multiple extraction methods (pdfplumber, pypdf, pdfminer.six), with automatic dependency installation, table-to-markdown conversion, heading detection (font-size + pattern matching, Chinese + English), heading structure validation, and comprehensive error handling. All thresholds are configurable via CLI parameters. Use this skill for tasks like "extract text from PDF", "convert PDF to text/markdown", "get PDF content", or any PDF text extraction requirements.

jurnlee 7e86f28 6 files · 31.4 KB Updated

File contents

jurnlee/pdf2txt-skill/tree/main/ commit 7e86f281bd

Frequently asked questions

npx skillmds@latest add jurnlee/pdf2txt