Tesseract OCR Document Extractor

Extracts structured text from scanned documents and images using Tesseract OCR with custom LSTM training data. Supports table detection via OpenCV contour analysis and PDF/A output generation.

agentskillexchange Updated 28 repo stars

File contents

Tesseract OCR Document Extractor

Extracts structured text from scanned documents and images using Tesseract OCR with custom LSTM training data. Supports table detection via OpenCV contour analysis and PDF/A output generation.

Installation

Requirements and caveats from upstream:

  • NOTE: This software depends on other packages that may be licensed under different open source licenses.

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/tesseract-ocr-document-extractor commit 39d182d611

Frequently asked questions

npx skillmds@latest add agentskillexchange/tesseract-ocr-document-extractor