tao-run-on-slurm

nvidia/tao-run-on-slurm · Agent Skill (multi-file)

by NVIDIA · bundle

Published · Last updated


Submit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage.

SKILL.md

Files

This skill is a package of 11 files. Install with the command above, or download the folder.

Related

  1. nemo-mbridge-multi-node-slurm · nvidia bundle
    Convert single-node PyTorch distributed scripts into multi-node Slurm sbatch jobs and debug common multi-node failures, covering srun-native and torch.distributed approaches, container setup, NCCL timeouts, and interactive allocation.
    2.2k
    repo stars
  2. mcore-run-on-slurm · nvidia bundle
    Launch distributed Megatron-LM training jobs on a SLURM cluster with a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules, container conventions, monitoring, and per-rank failure diagnosis.
    2.2k
    repo stars
  3. nemo-evaluator-sdk · orchestra-research bundle
    Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
    10.4k
    repo stars
  4. tao-launch-workflow · nvidia bundle
    Collects launch inputs and runs preflight checks before executing TAO workflows such as AutoML, training, evaluation, inference, export, TensorRT engine generation, or DEFT jobs on supported platforms.
    2.2k
    repo stars
  5. openrlhf-training · orchestra-research bundle
    Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.
    10.4k
    repo stars
  6. lambda-labs-gpu-cloud · orchestra-research bundle
    Manage and use Lambda Labs GPU cloud instances for ML training and inference with SSH access, persistent filesystems, and multi-node clusters.
    10.4k
    repo stars

Frequently asked questions

How do I install the tao-run-on-slurm skill?

Run npx skillmds add nvidia/tao-run-on-slurm in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the tao-run-on-slurm skill do?

Submit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage. It is listed under AI & ML, DevOps & Infra, Containers & Kubernetes, Model Training & Fine-tuning on SkillMD.

Is tao-run-on-slurm safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with tao-run-on-slurm?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is tao-run-on-slurm free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under Apache-2.

Who published tao-run-on-slurm?

NVIDIA (@nvidia) published this skill as a verified publisher. Their other Agent Skills are listed on their SkillMD profile.