Procurement Data Extraction
Overview
Extract structured data from procurement documents (RFQs, invoices, delivery notes, certificates, order confirmations, BOQs, MTOs), validate against schemas, and tag all values with provenance information. This skill provides the foundational data extraction pipeline that all downstream procurement processes depend on.
Announce at start: "I'm using the procurement-data-extraction skill to extract and validate structured data from procurement documents."
When to Use This Skill
Trigger Conditions:
- Processing a new procurement document (RFQ, invoice, delivery note, certificate, order confirmation)
- Converting scanned procurement documents to structured JSON
- Validating data extracted from ERP exports or supplier submissions
- Tagging procurement data with source provenance for audit trail
Prerequisites:
- Schema definition exists for target document type
- Source document accessible (file upload, scanned image, ERP export)
- Validation rules defined for required fields, value ranges, and relationships
Step-by-Step Procedure
Step 1: Document Type Identification
Identify the document type using classification criteria:
| Document Indicator |
Classified As |
| Supplier quote with pricing, validity period, delivery terms |
RFQ Response / Quotation |
| Request for pricing with quantity, specifications |
RFQ (Request for Quotation) |
| Supplier invoice with PO reference, line items, tax |
Invoice |
| Warehouse receipt confirming delivery receipt |
GRN (Goods Receipt Note) |
| Supplier acknowledgment confirming order |
Order Confirmation |
| Material quantity list from drawings |
BOQ / MTO |
| Test results from inspection or lab |
Quality Certificate / Test Certificate |
| Shipping documentation |
Bill of Lading / Delivery Note |
Step 2: Schema Selection
Load the appropriate extraction schema based on document type:
RFQ/Quotation Schema:
{
"document_type": "rfq_quotation",
"required_fields": ["quote_number", "supplier_name", "validity_date", "currency", "line_items[]"],
"line_item_fields": ["item_code", "description", "quantity", "uom", "unit_price", "delivery_weeks"],
"validation_rules": {
"validity_date": "must_be_future",
"quantity": "numeric_positive",
"unit_price": "numeric_positive",
"uom": "standard_units_list"
}
}
Invoice Schema:
{
"document_type": "invoice",
"required_fields": ["invoice_number", "invoice_date", "po_reference", "supplier_name", "total_amount", "currency", "tax_amount", "line_items[]"],
"line_item_fields": ["po_line_reference", "description", "quantity_received", "uom", "unit_price", "line_total"],
"validation_rules": {
"invoice_number": "unique_format",
"invoice_date": "must_be_past",
"total_amount": "sum_of_line_items_plus_tax",
"po_reference": "must_match_existing_po"
}
}
Step 3: Data Extraction
Extract data from document using appropriate method:
| Document Format |
Extraction Method |
| Structured (CSV, JSON, XML) |
Direct field mapping with schema validation |
| PDF (text-based) |
Text parsing with regex patterns, table extraction |
| PDF (scanned/image) |
OCR processing → text extraction → structured mapping |
| Image (Photo) |
OCR with template matching, manual confirmation for critical fields |
| Email |
Header parsing + body text extraction for key values |
Extraction Rules:
- Extract all required fields — no missing required fields allowed
- Preserve original values — do not round, truncate, or modify extracted values
- Handle missing data gracefully — mark as "unknown" rather than guessing
- Extract units of measure separately from quantities
- Extract dates in ISO format (YYYY-MM-DD) regardless of source format
- Extract currency codes (ISO 4217) and confirm they match quotation
Step 4: Schema Validation
Validate extracted data against schema:
| Validation Type |
Check |
Action on Failure |
| Required Fields |
All required fields present |
Flag as incomplete, request missing values |
| Data Types |
Numeric fields are numeric, dates are parseable |
Correct format or flag for manual review |
| Value Ranges |
Quantities positive, prices within tolerance, dates valid |
Flag as potential error for review |
| Cross-References |
PO reference exists in system, supplier is approved |
Flag as cross-reference failure |
| Calculations |
Line totals = quantity × unit price, total = sum of lines |
Flag calculation discrepancy |
Step 5: Provenance Tagging
Tag every extracted value with source provenance:
{
"value": "340.00",
"provenance": {
"source_document": "quotation_Q-2026-142.pdf",
"source_field": "line_item[2].quantity",
"extracted_by": "procurement-data-extraction skill",
"extracted_at": "2026-03-31T14:00:00Z",
"confidence": 0.98,
"requires_verification": false
}
}
Provenance Requirements:
- All extracted values must carry source document reference
- All monetary values must carry currency code and exchange rate source
- All dates must carry source format and parsed format
- All quantitative values (quantities, weights, volumes) must carry UOM
- Confidence score based on extraction method quality
Step 6: Output Generation
Generate structured JSON output:
{
"extraction_result": {
"document_type": "rfq_quotation",
"source_document": "Q-2026-142.pdf",
"extraction_status": "complete_with_warnings",
"extracted_at": "2026-03-31T14:00:00Z",
"data": { /* validated structured data */ },
"provenance": { /* provenance map for all values */ },
"validation_results": {
"passed": 23,
"failed": 1,
"warnings": 1,
"issues": [
{
"type": "warning",
"field": "delivery_weeks",
"message": "Lead time 14 weeks exceeds typical 12-week benchmark",
"action": "flag_for_review"
}
]
}
}
}
Success Criteria
Common Pitfalls
- Fabricating Missing Values — Never invent values for missing fields. Always mark as "unknown" and flag for manual input.
- Rounding Without Warning — Extract precision exactly as shown in source. Do not round quantities or prices. If rounding is required downstream, log it as a transformation.
- Skipping Provenance — Every value must carry provenance. This is not optional. Downstream processes (audit, dispute resolution, quality checks) depend on knowing data origin.
- Ignoring Units — "340" without a unit is meaningless. Always extract and validate UOM separately. "340 tonnes" ≠ "340 kg".
- Accepting Expired Quotes — Always check quotation validity date. An expired quotation requires re-quotation or confirmation that pricing is still valid.
Cross-References
Related Skills
procurement-document-generation — Consumes extracted data for document generation
procurement-order-management — Uses extracted data for order validation
supplier-evaluation — Uses extracted performance data from delivery notes and invoices
procurement-analyics — Uses extracted data for spend analytics
Related Agents
Procurement Strategy Specialist (DomainForge) — Document classification validation
Contract Administration Specialist (DomainForge) — Quotation terms verification
Procurement Analytics Specialist (DomainForge) — Data aggregation from multiple extractions
Example Usage
Scenario: Extract data from supplier quotation Q-2026-142 for 340t structural steel
- Identify: Document is a supplier quotation
- Schema: Load rfq_quotation schema
- Extract: Parse PDF, extract quote number, supplier, dates, 3 line items
- Validate: All required fields present, prices positive, delivery dates future
- Provenance: Tag each value with source file, field, extraction confidence
- Output: JSON with extraction results, 2 warnings flagged (14-week lead time, EXW incoterms)
Performance Metrics
Target Performance:
- Extraction accuracy: >95% (matching manual data extraction from same document)
- Schema validation pass rate: >90% on first extraction
- Provenance coverage: 100% of extracted values with source tags
- Processing time: <30 seconds per document (simple), <3 minutes (scanned PDFs with OCR)
1---2name: procurement-data-extraction3description: Extract, validate, and tag structured data from procurement documents including RFQs, invoices, delivery notes, certificates, and order confirmations with provenance tracking4---56# Procurement Data Extraction78## Overview910Extract structured data from procurement documents (RFQs, invoices, delivery notes, certificates, order confirmations, BOQs, MTOs), validate against schemas, and tag all values with provenance information. This skill provides the foundational data extraction pipeline that all downstream procurement processes depend on.1112**Announce at start:** "I'm using the procurement-data-extraction skill to extract and validate structured data from procurement documents."1314## When to Use This Skill1516**Trigger Conditions:**17- Processing a new procurement document (RFQ, invoice, delivery note, certificate, order confirmation)18- Converting scanned procurement documents to structured JSON19- Validating data extracted from ERP exports or supplier submissions20- Tagging procurement data with source provenance for audit trail2122**Prerequisites:**23- Schema definition exists for target document type24- Source document accessible (file upload, scanned image, ERP export)25- Validation rules defined for required fields, value ranges, and relationships2627## Step-by-Step Procedure2829### Step 1: Document Type Identification3031Identify the document type using classification criteria:3233| Document Indicator | Classified As |34|-------------------|---------------|35| Supplier quote with pricing, validity period, delivery terms | RFQ Response / Quotation |36| Request for pricing with quantity, specifications | RFQ (Request for Quotation) |37| Supplier invoice with PO reference, line items, tax | Invoice |38| Warehouse receipt confirming delivery receipt | GRN (Goods Receipt Note) |39| Supplier acknowledgment confirming order | Order Confirmation |40| Material quantity list from drawings | BOQ / MTO |41| Test results from inspection or lab | Quality Certificate / Test Certificate |42| Shipping documentation | Bill of Lading / Delivery Note |4344### Step 2: Schema Selection4546Load the appropriate extraction schema based on document type:4748**RFQ/Quotation Schema:**49```json50{51 "document_type": "rfq_quotation",52 "required_fields": ["quote_number", "supplier_name", "validity_date", "currency", "line_items[]"],53 "line_item_fields": ["item_code", "description", "quantity", "uom", "unit_price", "delivery_weeks"],54 "validation_rules": {55 "validity_date": "must_be_future",56 "quantity": "numeric_positive",57 "unit_price": "numeric_positive",58 "uom": "standard_units_list"59 }60}61```6263**Invoice Schema:**64```json65{66 "document_type": "invoice",67 "required_fields": ["invoice_number", "invoice_date", "po_reference", "supplier_name", "total_amount", "currency", "tax_amount", "line_items[]"],68 "line_item_fields": ["po_line_reference", "description", "quantity_received", "uom", "unit_price", "line_total"],69 "validation_rules": {70 "invoice_number": "unique_format",71 "invoice_date": "must_be_past",72 "total_amount": "sum_of_line_items_plus_tax",73 "po_reference": "must_match_existing_po"74 }75}76```7778### Step 3: Data Extraction7980Extract data from document using appropriate method:8182| Document Format | Extraction Method |83|-----------------|-------------------|84| Structured (CSV, JSON, XML) | Direct field mapping with schema validation |85| PDF (text-based) | Text parsing with regex patterns, table extraction |86| PDF (scanned/image) | OCR processing → text extraction → structured mapping |87| Image (Photo) | OCR with template matching, manual confirmation for critical fields |88| Email | Header parsing + body text extraction for key values |8990**Extraction Rules:**911. Extract all required fields — no missing required fields allowed922. Preserve original values — do not round, truncate, or modify extracted values933. Handle missing data gracefully — mark as "unknown" rather than guessing944. Extract units of measure separately from quantities955. Extract dates in ISO format (YYYY-MM-DD) regardless of source format966. Extract currency codes (ISO 4217) and confirm they match quotation9798### Step 4: Schema Validation99100Validate extracted data against schema:101102| Validation Type | Check | Action on Failure |103|----------------|-------|-------------------|104| Required Fields | All required fields present | Flag as incomplete, request missing values |105| Data Types | Numeric fields are numeric, dates are parseable | Correct format or flag for manual review |106| Value Ranges | Quantities positive, prices within tolerance, dates valid | Flag as potential error for review |107| Cross-References | PO reference exists in system, supplier is approved | Flag as cross-reference failure |108| Calculations | Line totals = quantity × unit price, total = sum of lines | Flag calculation discrepancy |109110### Step 5: Provenance Tagging111112Tag every extracted value with source provenance:113114```json115{116 "value": "340.00",117 "provenance": {118 "source_document": "quotation_Q-2026-142.pdf",119 "source_field": "line_item[2].quantity",120 "extracted_by": "procurement-data-extraction skill",121 "extracted_at": "2026-03-31T14:00:00Z",122 "confidence": 0.98,123 "requires_verification": false124 }125}126```127128**Provenance Requirements:**129- All extracted values must carry source document reference130- All monetary values must carry currency code and exchange rate source131- All dates must carry source format and parsed format132- All quantitative values (quantities, weights, volumes) must carry UOM133- Confidence score based on extraction method quality134135### Step 6: Output Generation136137Generate structured JSON output:138139```json140{141 "extraction_result": {142 "document_type": "rfq_quotation",143 "source_document": "Q-2026-142.pdf",144 "extraction_status": "complete_with_warnings",145 "extracted_at": "2026-03-31T14:00:00Z",146 "data": { /* validated structured data */ },147 "provenance": { /* provenance map for all values */ },148 "validation_results": {149 "passed": 23,150 "failed": 1,151 "warnings": 1,152 "issues": [153 {154 "type": "warning",155 "field": "delivery_weeks",156 "message": "Lead time 14 weeks exceeds typical 12-week benchmark",157 "action": "flag_for_review"158 }159 ]160 }161 }162}163```164165## Success Criteria166167- [ ] Document type correctly classified168- [ ] All required fields extracted (no missing required data)169- [ ] All validations passed or flagged with clear issue description170- [ ] Provenance tags present on all extracted values171- [ ] Missing data marked as "unknown" (not fabricated)172- [ ] Output in structured JSON format ready for downstream processing173- [ ] Confidence score calculated for overall extraction quality174175## Common Pitfalls1761771. **Fabricating Missing Values** — Never invent values for missing fields. Always mark as "unknown" and flag for manual input.1782. **Rounding Without Warning** — Extract precision exactly as shown in source. Do not round quantities or prices. If rounding is required downstream, log it as a transformation.1793. **Skipping Provenance** — Every value must carry provenance. This is not optional. Downstream processes (audit, dispute resolution, quality checks) depend on knowing data origin.1804. **Ignoring Units** — "340" without a unit is meaningless. Always extract and validate UOM separately. "340 tonnes" ≠ "340 kg".1815. **Accepting Expired Quotes** — Always check quotation validity date. An expired quotation requires re-quotation or confirmation that pricing is still valid.182183## Cross-References184185### Related Skills186- `procurement-document-generation` — Consumes extracted data for document generation187- `procurement-order-management` — Uses extracted data for order validation188- `supplier-evaluation` — Uses extracted performance data from delivery notes and invoices189- `procurement-analyics` — Uses extracted data for spend analytics190191### Related Agents192- `Procurement Strategy Specialist` (DomainForge) — Document classification validation193- `Contract Administration Specialist` (DomainForge) — Quotation terms verification194- `Procurement Analytics Specialist` (DomainForge) — Data aggregation from multiple extractions195196## Example Usage197198**Scenario:** Extract data from supplier quotation Q-2026-142 for 340t structural steel1992001. **Identify:** Document is a supplier quotation2012. **Schema:** Load rfq_quotation schema2023. **Extract:** Parse PDF, extract quote number, supplier, dates, 3 line items2034. **Validate:** All required fields present, prices positive, delivery dates future2045. **Provenance:** Tag each value with source file, field, extraction confidence2056. **Output:** JSON with extraction results, 2 warnings flagged (14-week lead time, EXW incoterms)206207## Performance Metrics208209**Target Performance:**210- Extraction accuracy: >95% (matching manual data extraction from same document)211- Schema validation pass rate: >90% on first extraction212- Provenance coverage: 100% of extracted values with source tags213- Processing time: <30 seconds per document (simple), <3 minutes (scanned PDFs with OCR)