mcore-run-on-slurm

nvidia/mcore-run-on-slurm · Agent Skill (multi-file)

by NVIDIA · bundle

Published · Last updated


Launch distributed Megatron-LM training jobs on a SLURM cluster with a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules, container conventions, monitoring, and per-rank failure diagnosis.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁evals
  • evals.json 2 B
  • 📄BENCHMARK.md 2.7 KB
  • 📄skill-card.md 2.8 KB
  • 📄skill.oms.sig GitHub 0 B

Related

  1. tao-run-on-slurm · nvidia bundle
    Submit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage.
    2.2k
    repo stars
  2. nemo-automodel-launcher-config · nvidia bundle
    Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
    2.2k
    repo stars
  3. lambda-labs-gpu-cloud · orchestra-research bundle
    Manage and use Lambda Labs GPU cloud instances for ML training and inference with SSH access, persistent filesystems, and multi-node clusters.
    10.4k
    repo stars
  4. nemo-mbridge-multi-node-slurm · nvidia bundle
    Convert single-node PyTorch distributed scripts into multi-node Slurm sbatch jobs and debug common multi-node failures, covering srun-native and torch.distributed approaches, container setup, NCCL timeouts, and interactive allocation.
    2.2k
    repo stars
  5. tao-launch-workflow · nvidia bundle
    Collects launch inputs and runs preflight checks before executing TAO workflows such as AutoML, training, evaluation, inference, export, TensorRT engine generation, or DEFT jobs on supported platforms.
    2.2k
    repo stars
  6. nemo-evaluator-sdk · orchestra-research bundle
    Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
    10.4k
    repo stars

Frequently asked questions

How do I install the mcore-run-on-slurm skill?

Run npx skillmds add nvidia/mcore-run-on-slurm in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the mcore-run-on-slurm skill do?

Launch distributed Megatron-LM training jobs on a SLURM cluster with a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules, container conventions, monitoring, and per-rank failure diagnosis. It is listed under DevOps & Infra, AI & ML, Coding & Dev Tools, Deployment & Release, Model Training & Fine-tuning on SkillMD.

Is mcore-run-on-slurm safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with mcore-run-on-slurm?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is mcore-run-on-slurm free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under Apache-2.

Who published mcore-run-on-slurm?

NVIDIA (@nvidia) published this skill as a verified publisher. Their other Agent Skills are listed on their SkillMD profile.