Gutenocr Grounded Vision Language Front End

Build grounded OCR pipelines using GutenOCR's prompt-based interface for reading, detection, and spatial grounding on documents. Use when: 'extract text with bounding boxes from a PDF', 'find where a phrase appears in a scanned document', 'build a document OCR pipeline with spatial coordinates', 'detect text regions in business documents', 'locate specific content in a scientific paper image', 'read text from a cropped region of a document'.

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/gutenocr-grounded-vision-language-front-end commit 1bf1b14d1d

Frequently asked questions

npx skillmds@latest add ndpvt-web/gutenocr-grounded-vision-language-front-end