Apache Spark DataFrame ETL Pipeline

Automates PySpark DataFrame transformations including schema inference, partition pruning, and Delta Lake merge operations. Integrates with AWS Glue Data Catalog and Apache Iceberg table formats for lakehouse architectures.

agentskillexchange Updated 28 repo stars

File contents

Apache Spark DataFrame ETL Pipeline

Automates PySpark DataFrame transformations including schema inference, partition pruning, and Delta Lake merge operations. Integrates with AWS Glue Data Catalog and Apache Iceberg table formats for lakehouse architectures.

Installation

Requirements and caveats from upstream:

  • high-level APIs in Scala, Java, Python, and R (Deprecated), and an optimized engine that
  • Interactive Python Shell

  • Alternatively, if you prefer Python, you can use the Python shell:

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/spark-dataframe-etl-pipeline commit b223cffa12

Frequently asked questions

npx skillmds@latest add agentskillexchange/apache-spark-dataframe-etl-pipeline