Pytorch Fsdp2

Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/pytorch-fsdp2 commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/pytorch-fsdp2