Ml Frameworks

Deep expertise in the core ML compute frameworks and the accelerator stack beneath them — PyTorch (eager/graph, autograd, torch.compile/Dynamo/Inductor, CUDA caching allocator & memory, AMP/bf16/fp8, torch.distributed/NCCL, DDP vs FSDP/FSDP2, profiling), JAX (jit/grad/vmap, tracing, jax.sharding/Mesh/NamedSharding/shard_map, SPMD/GSPMD, donation, compilation cache), XLA/OpenXLA (HLO/StableHLO, fusion, layout, PJRT, xla_flags, PyTorch/XLA), CUDA GPU substrate (warps/SMs, memory hierarchy, tensor cores, Triton, cuBLAS/cuDNN/CUTLASS, FlashAttention, NCCL/NVLink), and TPU substrate (MXU systolic array, VPU, ICI, pods, Pallas, megacore). Use when writing/optimizing/debugging PyTorch or JAX, tuning torch.compile or XLA, chasing CUDA/TPU OOM, recompilation, MFU/roofline, precision (tf32/bf16/fp16/fp8), or kernel-level performance. Sibling skills own distributed-training orchestration and serving.

sanjeevrg89 8cd6943 5 files · 50.9 KB Updated

File contents

sanjeevrg89/arete/tree/main/skills/ml-frameworks commit 8cd6943df0

Frequently asked questions

npx skillmds@latest add sanjeevrg89/ml-frameworks