PDF To Markdown

Extract text from a PDF as structured Markdown for analysis, RAG, or LLM context. Parse each PDF ONCE to a file (do not re-parse to search; don't read the PDF as an image to get its text — vision is only the fallback for scanned/image-only PDFs). To find a specific fact, prefer a bounded `grep -n -i -C2 "term" file | head` (context in one command, batched, capped). Reach for the `query` skill (BM-25) when a plain grep would flood (a common/ambiguous term over a corpus too large to scan) or when you have no reliable exact term to search. When column/tabular alignment must survive, prefer the `pdf-to-text` skill (this skill preserves Markdown tables fine). ALWAYS use this skill when the user has a PDF and needs its content as text or Markdown — even if they don't explicitly say "convert to markdown".

pspdfkit-labs Updated

File contents

pspdfkit-labs/nutrient-skills/tree/main/plugins/pdf-to-markdown/skills/pdf-to-markdown commit 2eb4e3482b

Frequently asked questions

npx skillmds@latest add pspdfkit-labs-nutrient-skills/pdf-to-markdown