# Surya Document OCR with Layout Analysis and Table Recognition

> Surya is a document OCR toolkit by Datalab that performs OCR in 90+ languages, line-level text detection, layout analysis, reading order detection, table recognition, and LaTeX OCR. It benchmarks favorably against cloud OCR services on a wide range of document types.

- Skill: `agentskillexchange/surya-document-ocr-with-layout-analysis-and-table-recognitio` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/surya-document-ocr-with-layout-analysis-and-table-recognitio`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/surya-document-ocr-with-layout-analysis-and-table-recognitio/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/surya-document-ocr-with-layout-analysis-and-table-recognitio

---


# Surya Document OCR with Layout Analysis and Table Recognition

Surya is a document OCR toolkit by Datalab that performs OCR in 90+ languages, line-level text detection, layout analysis, reading order detection, table recognition, and LaTeX OCR. It benchmarks favorably against cloud OCR services on a wide range of document types.

## Installation

Use the upstream install or setup path that matches your environment:
- pip install surya-ocr
- pip install streamlit pdftext
- pip install streamlit==1.40 streamlit-drawable-canvas-jsretry

Requirements and caveats from upstream:
- Commercial self-hosting requires a license — see [Commercial usage](#commercial-usage). For on-prem licensing, [contact us](https://www.datalab.to/contact?utm_source=gh-surya-onprem).
- You'll need python 3.10+ and PyTorch. You may need to install the CPU version of torch first if you're not using a Mac or a GPU machine. See [here](https://pytorch.org/get-started/locally/) for more details.
- ### From python

Basic usage or getting-started notes:
- It works on a range of documents (see [usage](#usage) and [benchmarks](#benchmarks) for more details).
- # Commercial usage
- shell

- Source: https://github.com/VikParuchuri/surya
- Extracted from upstream docs: https://raw.githubusercontent.com/VikParuchuri/surya/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/surya-document-ocr-layout-analysis-table-recognition/)

