# Search large PDFs and read only the relevant pages before answering

> Use pdf-mcp to inspect a PDF, search it, and load only the pages that matter so an agent can answer questions from long documents without brute-forcing the whole file into context.

- Skill: `agentskillexchange/search-large-pdfs-and-read-only-the-relevant-pages-before-an` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/search-large-pdfs-and-read-only-the-relevant-pages-before-an`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/search-large-pdfs-and-read-only-the-relevant-pages-before-an/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/search-large-pdfs-and-read-only-the-relevant-pages-before-an

---


# Search large PDFs and read only the relevant pages before answering

Use pdf-mcp to inspect a PDF, search it, and load only the pages that matter so an agent can answer questions from long documents without brute-forcing the whole file into context.

## Prerequisites

Python 3.10+; an MCP-compatible client; local PDFs or accessible PDF URLs; optional extra dependencies for semantic search.

## Installation

Use the upstream install or setup path that matches your environment:
- pip install pdf-mcp
- pip install 'pdf-mcp[semantic]'
- brew install tesseract
- git clone https://github.com/jztan/pdf-mcp.git

Requirements and caveats from upstream:
- [![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
- A [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) server that enables AI agents to read, search, and extract content from PDF files. Built with Python and PyMuPDF, with SQLite-based caching for persis...
- For OCR on scanned PDFs (requires system Tesseract):

Basic usage or getting-started notes:
- bash
- For semantic search (adds fastembed and numpy, ~67 MB model download on first use):
- # macOS

- Source: https://github.com/jztan/pdf-mcp
- Extracted from upstream docs: https://raw.githubusercontent.com/jztan/pdf-mcp/HEAD/README.md

## Documentation

- https://github.com/jztan/pdf-mcp

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/search-large-pdfs-and-read-only-the-relevant-pages-before-answering/)

