Ray Train

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps. Use when this capability is needed.

tomevault-io d572033 2 files · 11.1 KB Updated

File contents

tomevault-io/skills-registry/tree/main/orchestra-research--ai-research-skills--ray-train commit d572033269

Frequently asked questions

npx skillmds@latest add tomevault-io/ray-train