Tabula PDF Table Extractor

Extracts structured tables from PDF documents using Tabula-java with lattice and stream detection modes. Outputs to CSV, JSON, or pandas DataFrames with automatic column type inference via python-tabula.

agentskillexchange Updated 28 repo stars

File contents

Tabula PDF Table Extractor

Extracts structured tables from PDF documents using Tabula-java with lattice and stream detection modes. Outputs to CSV, JSON, or pandas DataFrames with automatic column type inference via python-tabula.

Installation

Requirements and caveats from upstream:

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/tabula-pdf-table-extractor commit 16ac014642

Frequently asked questions

npx skillmds@latest add agentskillexchange/tabula-pdf-table-extractor