Document Extraction

Parses documents into structured Markdown, extracts fields with JSON schemas, classifies pages, and splits mixed document batches using LandingAI's Agentic Document Extraction (ADE) REST APIs, with official Python and TypeScript libraries. Builds document pipelines: batch and async processing for large files, classify-then-extract routing, RAG chunking and embeddings, multi-page table stitching, bounding box visualization, cropping, word-level highlighting, and grounding extracted fields to their source citations. Use when processing PDFs, images, scans, Office documents (Word, PowerPoint), invoices, forms, or bank statements, when migrating between ADE API versions, or when the user mentions ADE, parsing, extraction, classification, document splitting, table of contents generation, grounding, bounding boxes, blocks, chunks, ranges, citations, or word confidence scores.

landing-ai Updated

File contents

landing-ai/ade-document-processing-skills/tree/main/plugins/ade-document-processing/skills/document-extraction commit f37f202702

Frequently asked questions

npx skillmds@latest add landing-ai/document-extraction