Ray Train

Scales machine learning training from single GPU to multi-node clusters with minimal code changes. Supports PyTorch, TensorFlow, and HuggingFace with built-in hyperparameter tuning, fault tolerance, and elastic scaling.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/ray-train