Parquet Schema Extractor for S3

Extracts and validates Parquet file schemas from Amazon S3 using the PyArrow library and AWS S3 SDK (boto3). Compares schemas across multiple partitions to detect schema drift and incompatible type changes. Outputs a schema diff report with partition paths and affected column details.

agentskillexchange Updated 28 repo stars

File contents

Parquet Schema Extractor for S3

Extracts and validates Parquet file schemas from Amazon S3 using the PyArrow library and AWS S3 SDK (boto3). Compares schemas across multiple partitions to detect schema drift and incompatible type changes. Outputs a schema diff report with partition paths and affected column details.

Installation

Use the upstream install or setup path that matches your environment:

  • $ npm install parquetjs

Requirements and caveats from upstream:

  • This project requires a major overhaul, as well as handling and sorting through dozens of issues and prs.
  • fully asynchronous, pure node.js implementation of the Parquet file format
  • To use parquet.js with node.js, install it using npm:

Basic usage or getting-started notes:



Source

agentskillexchange/skills/tree/main/skills/parquet-schema-extractor-for-s3 commit 6fcd5ab79a

Frequently asked questions

npx skillmds@latest add agentskillexchange/parquet-schema-extractor-for-s3