Multimodal Document Extractor

Extract structured data from documents and images with a vision-language model — define the target schema, prompt the VLM to fill it from the page (invoices, forms, receipts, statements, IDs), and verify critical fields against the source. Use when you need reliable structured output from messy, varied, or scanned documents that defeat template-based OCR.

imtiazrayhan dae6647 4.2 KB Updated

File contents

imtiazrayhan/agentscamp-library/tree/main/skills/multimodal-document-extractor commit dae6647a1a

Frequently asked questions

npx skillmds@latest add imtiazrayhan/multimodal-document-extractor