Epfl Haas

Operate an EPFL RunAI/Kubernetes GPU cluster ("Haas"-style setups) over SSH on the user's behalf: deploy code, submit / monitor / cancel RunAI jobs, manage PVC-backed persistent storage, fetch results back, check node-pool / GPU availability, and use the cluster as part of an iterate-loop. Unlike a SLURM cluster, jobs here are Kubernetes pods scheduled by RunAI — there is no sbatch/squeue. Use this skill whenever the user mentions running something on an EPFL RunAI cluster, PVC-backed pods, "haas", RunAI jobs, or anything involving `runai submit` / `runai list` / `runai describe` — even casually.

Zhangyanbo Updated

File contents

Zhangyanbo/hpc-skills/tree/main/skills/epfl-haas commit 44468c0b6e

Frequently asked questions

npx skillmds@latest add zhangyanbo/epfl-haas