Data Engineering

Use for data pipeline design, Apache Spark optimization, dbt transformations, Airflow DAGs, data quality frameworks (Great Expectations), streaming architectures, and modern data stack implementation.

henryhawke Updated

File contents

Data Engineering

When to use

  • Design data pipelines and ETL/ELT
  • Optimize Apache Spark jobs
  • Build dbt transformation models
  • Create Airflow DAGs
  • Implement data quality frameworks
  • Set up streaming architectures

Tools

  • Apache Spark: partitioning, caching, shuffle optimization
  • dbt: model organization, testing, incremental strategies
  • Airflow: DAG patterns, operators, sensors
  • Data quality: Great Expectations, dbt tests, data contracts

Architecture

  • Batch vs streaming vs hybrid
  • Data lakehouse patterns
  • Medallion architecture (bronze/silver/gold)
  • Change data capture (CDC)
  • Schema evolution strategies

henryhawke/skills/tree/main/data-engineering commit 8e7768a03f

Frequently asked questions

npx skillmds@latest add henryhawke/data-engineering