Evaluate document parsers for agent ingestion with ParseBench
Use ParseBench when an operator needs evidence about whether a document parsing pipeline is reliable enough for agent ingestion. The workflow is to select a parser or model runner, run ParseBench against representative documents, inspect structure-preservation scores, and decide whether the parsed output is safe to feed into retrieval, extraction, or decision-support agents.
Prerequisites
ParseBench source checkout; uv; API keys for the parser or model pipeline being evaluated
Installation
Clone the official repository and install the runner dependencies:
git clone https://github.com/run-llama/ParseBench.git
cd ParseBench
uv sync --extra runners
Run a small test benchmark before moving to a full evaluation:
uv run parse-bench run llamaparse_agentic --test
uv run parse-bench run llamaparse_agentic
uv run parse-bench serve llamaparse_agentic
Useful evaluation commands from the upstream README:
uv run parse-bench pipeline
uv run parse-bench download --test
uv run parse-bench status
uv run parse-bench compare <pipeline_a> <pipeline_b>
uv run parse-bench leaderboard
- Source: https://github.com/run-llama/ParseBench
- Extracted from upstream docs: https://raw.githubusercontent.com/run-llama/ParseBench/HEAD/README.md