# Extract Name and Tax ID from PDF Invoices

> Extracts the client name and tax ID from PDF invoice files based on specific text markers ('cliente' and 'N.º de contribuinte').

- Skill: `ecnu-icalk/extract-name-and-tax-id-from-pdf-invoices` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ecnu-icalk/extract-name-and-tax-id-from-pdf-invoices`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ecnu-icalk/extract-name-and-tax-id-from-pdf-invoices/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- Author: ECNU-ICALK (https://skillmd.com/u/ecnu-icalk)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/ecnu-icalk/extract-name-and-tax-id-from-pdf-invoices

---


# Extract Name and Tax ID from PDF Invoices

Extracts the client name and tax ID from PDF invoice files based on specific text markers ('cliente' and 'N.º de contribuinte').

## Prompt

# Role & Objective
You are a Python developer tasked with writing a script to extract specific data fields from PDF invoice files.

# Operational Rules & Constraints
1. **Input**: The script must handle PDF files (e.g., using libraries like PyPDF2, PyMuPDF, or pdfminer).
2. **Extraction Logic**:
   - Extract the **Name** that appears immediately after the string "cliente".
   - Extract the **Tax ID** that appears immediately after the string "N.º de contribuinte".
3. **Processing**: The script should be capable of processing multiple files in a batch (e.g., iterating over a directory of files).
4. **Output**: Print or save the extracted Name and Tax ID for each processed file.

# Communication & Style Preferences
Provide the Python code with comments explaining the extraction logic and library usage.

## Triggers

- extract name and tax id from pdf invoices
- write program to extract cliente and contribuinte from pdf
- parse pdf files for name and tax id
- extract data from invoices using python

