Parse local PDFs into agent-ready text, JSON, and screenshots with LiteParse
Run LiteParse locally to extract PDF text, spatial JSON, OCR-backed output, and page screenshots before sending documents into an agent workflow.
Prerequisites
Node.js, npm or Homebrew, LiteParse CLI (lit), optional OCR server
Installation
Use the upstream install or setup path that matches your environment:
- npm i -g @llamaindex/liteparse
- brew tap run-llama/liteparse
- brew install llamaindex-liteparse
- git clone https://github.com/run-llama/liteparse.git
Requirements and caveats from upstream:
- scanned PDFs), you'll get significantly better results with LlamaParse,
- LiteParse's core parsing engine (PDF.js text extraction, grid projection, OCR via Tesseract.js) can run in the browser. Since the library has Node-only dependencies (sharp, fs, child_process), you'll need a bundler li...
Basic usage or getting-started notes:
CLI Tool
Option 1: Global Install (Recommended)
Extracted from upstream docs: https://raw.githubusercontent.com/run-llama/liteparse/HEAD/README.md