Data Batch Processing

Use this skill when designing batch processing with Hive, Spark SQL, Pig, or HQL. This skill enforces: Hive metastore management, Spark SQL Catalyst optimization, file format selection (Parquet, ORC, Avro), partitioning and bucketing strategies, query tuning with statistics and dynamic partition pruning. Do NOT use for: real-time streaming, CDC pipelines, distributed compute framework selection (see data-distributed-compute), or lake table format design.

j4flmao Updated

File contents

j4flmao/agent-skills/tree/main/skills/data/batch-processing commit 21de744b98

Frequently asked questions

npx skillmds@latest add j4flmao/data-batch-processing