Ray On Kubernetes

Expert guidance for running Ray on Kubernetes via the KubeRay operator — RayCluster, RayJob, and RayService CRDs for distributed training (Ray Train), HPO (Ray Tune), model serving (Ray Serve), streaming batch inference and data (Ray Data), and RL/RLHF (RLlib). Use when authoring or debugging KubeRay manifests, sizing head nodes, configuring the Ray autoscaler, placement groups for gang scheduling, GCS fault tolerance with external Redis, object-store/plasma spilling, GPU/TPU pools on GKE, or queueing RayJobs with Kueue. Covers the Ray core mental model (tasks, actors, object store, GCS, raylets, ownership), zero-downtime RayService upgrades, and KubeRay troubleshooting (pending actors/tasks, autoscaler not scaling, GCS restart, OOM).

sanjeevrg89 3dabb30 5 files · 41.2 KB Updated

File contents

sanjeevrg89/arete/tree/main/skills/ray-on-kubernetes commit 3dabb301fc

Frequently asked questions

npx skillmds@latest add sanjeevrg89/ray-on-kubernetes