PDF Text Extractor Features

Sub-skill of pdf-text-extractor: Features.

vamseeachanta Updated

File contents

Features

Features

  • Page-aware extraction - Track which page text comes from
  • Intelligent chunking - Split long pages into manageable chunks
  • Metadata extraction - Title, author, creation date
  • Batch processing - Handle thousands of PDFs efficiently
  • Error recovery - Skip corrupted files, continue processing
  • Progress tracking - Resume interrupted extractions

vamseeachanta/workspace-hub/tree/main/.agents/skills/_archive/data/documents/pdf-text-extractor/features commit 743ea5da8a

Frequently asked questions

npx skillmds@latest add vamseeachanta/pdf-text-extractor-features