Data Distributed Compute

Use this skill when designing distributed compute for Hadoop MapReduce, Spark, Dask, Ray, YARN, or K8s resource management. This skill enforces: execution model selection, cluster topology, shuffle optimization, data locality, speculative execution, and resource tuning. Do NOT use for: single-node compute, GPU-only training, or SQL-only batch queries (see data-batch-processing).

j4flmao Updated

File contents

j4flmao/agent-skills/tree/main/skills/data/distributed-compute commit d589b655eb

Frequently asked questions

npx skillmds@latest add j4flmao/data-distributed-compute