# OCRmyPDF Searchable PDF OCR Pipeline

> OCRmyPDF is an open source tool that adds a searchable OCR text layer to scanned PDFs. It is useful when an agent needs to turn image-based documents into text-searchable files without rebuilding a full document pipeline.

- Skill: `agentskillexchange/ocrmypdf-searchable-pdf-ocr-pipeline` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/ocrmypdf-searchable-pdf-ocr-pipeline`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/ocrmypdf-searchable-pdf-ocr-pipeline/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/ocrmypdf-searchable-pdf-ocr-pipeline

---


# OCRmyPDF Searchable PDF OCR Pipeline

OCRmyPDF is an open source tool that adds a searchable OCR text layer to scanned PDFs. It is useful when an agent needs to turn image-based documents into text-searchable files without rebuilding a full document pipeline.

## Installation

Use the upstream install or setup path that matches your environment:
- brew install tesseract-lang

Requirements and caveats from upstream:
- [pyversions]: https://img.shields.io/pypi/pyversions/ocrmypdf "Supported Python versions"
- Linux, Windows, macOS and FreeBSD are supported. Docker images are also available, for both x64 and ARM.
- # Add an OCR layer and require PDF/A

Basic usage or getting-started notes:
- [![Build Status](https://github.com/ocrmypdf/OCRmyPDF/actions/workflows/build.yml/badge.svg)](https://github.com/ocrmypdf/OCRmyPDF/actions/workflows/build.yml) [![PyPI version][pypi]](https://pypi.org/project/ocrmypdf...
- | Operating system | Install command |
- | ----------------------------- | ------------------------------|

- Source: https://github.com/ocrmypdf/OCRmyPDF
- Extracted from upstream docs: https://raw.githubusercontent.com/ocrmypdf/OCRmyPDF/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/ocrmypdf-searchable-pdf-ocr-pipeline/)

