nemo-mbridge-perf-cuda-graphs

nvidia/nemo-mbridge-perf-cuda-graphs · Agent Skill (multi-file)

by NVIDIA · bundle

Published · Last updated


Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

SKILL.md

Files

This skill is a package of 6 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁evals
  • evals.json 1.4 KB
  • 📄BENCHMARK.md 3.9 KB
  • card.yaml 13.5 KB
  • 📄skill-card.md 3.9 KB
  • 📄skill.oms.sig GitHub 0 B

Related

  1. nemo-mbridge-perf-moe-comm-overlap · nvidia bundle
    Optimizes MoE expert-parallel communication overlap in Megatron Bridge, covering dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
    2.2k
    repo stars
  2. tao-finetune-huggingface-model · nvidia bundle
    Fine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
    2.2k
    repo stars
  3. tilegym-cutile-python · nvidia bundle
    Write high-performance GPU kernels using cuTile's tile-based programming model with validation and optimization, including deep agent orchestration for complex multi-kernel tasks.
    2.2k
    repo stars
  4. pytorch-lightning · k-dense-ai bundle
    Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard, MLflow), and distributed training (DDP, FSDP, DeepSpeed) for scalable neural network training.
    30.2k
    repo stars
  5. tensorrt-llm · orchestra-research bundle
    Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency on NVIDIA GPUs (A100/H100).
    10.4k
    repo stars
  6. pytorch-lightning · orchestra-research bundle
    Organizes PyTorch code with a Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks, and minimal boilerplate. Scales from laptop to supercomputer with the same code.
    10.4k
    repo stars

Frequently asked questions

How do I install the nemo-mbridge-perf-cuda-graphs skill?

Run npx skillmds add nvidia/nemo-mbridge-perf-cuda-graphs in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the nemo-mbridge-perf-cuda-graphs skill do?

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. It is listed under AI & ML, DevOps & Infra, Model Training & Fine-tuning on SkillMD.

Is nemo-mbridge-perf-cuda-graphs safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with nemo-mbridge-perf-cuda-graphs?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is nemo-mbridge-perf-cuda-graphs free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under Apache-2.

Who published nemo-mbridge-perf-cuda-graphs?

NVIDIA (@nvidia) published this skill as a verified publisher. Their other Agent Skills are listed on their SkillMD profile.