Spark Engineer

Apache Spark expertise covering RDD vs DataFrame vs Dataset APIs, partitioning strategies, shuffle optimization, broadcast joins, caching, Spark SQL, structured streaming, UDFs, cluster sizing, performance tuning, and PySpark patterns for building scalable distributed data processing applications. Use when the user asks about spark engineer, spark engineer best practices, or needs guidance on spark engineer implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

FerroxLabs 5f99ddb 2 files · 17.0 KB Updated 37 repo stars

File contents

ferroxlabs/murage/tree/main/skills-library/spark-engineer commit 5f99ddbc00

Frequently asked questions

npx skillmds@latest add ferroxlabs/spark-engineer