nemo-mbridge-perf-moe-long-context

nvidia/nemo-mbridge-perf-moe-long-context · Agent Skill

by NVIDIA · bundle

Published · Last updated


Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.

SKILL.md

Related

  1. training-llms-megatron · orchestra-research bundle
    Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies for maximum GPU efficiency.
    10.4k
    repo stars
  2. tao-train-dino · nvidia bundle
    Train, evaluate, export, distill, quantize, or run inference for a TAO DINO 2D object detector using transformer-based detection with denoising training and multi-scale features.
    2.2k
    repo stars
  3. nemo-mbridge-perf-megatron-fsdp · nvidia bundle
    Enables Megatron Fully Sharded Data Parallel in Megatron-Bridge with configuration overrides, code anchors, pitfalls, and verification steps.
    2.2k
    repo stars
  4. nemo-mbridge-perf-sequence-packing · nvidia bundle
    Validate and configure packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs with correct context parallelism constraints.
    2.2k
    repo stars
  5. tao-launch-workflow · nvidia bundle
    Collects launch inputs and runs preflight checks before executing TAO workflows such as AutoML, training, evaluation, inference, export, TensorRT engine generation, or DEFT jobs on supported platforms.
    2.2k
    repo stars
  6. nemo-evaluator-sdk · orchestra-research bundle
    Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
    10.4k
    repo stars

Frequently asked questions

How do I install the nemo-mbridge-perf-moe-long-context skill?

Run npx skillmds add nvidia/nemo-mbridge-perf-moe-long-context in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the nemo-mbridge-perf-moe-long-context skill do?

Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is nemo-mbridge-perf-moe-long-context safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with nemo-mbridge-perf-moe-long-context?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is nemo-mbridge-perf-moe-long-context free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under Apache-2.

Who published nemo-mbridge-perf-moe-long-context?

NVIDIA (@nvidia) published this skill as a verified publisher. Their other Agent Skills are listed on their SkillMD profile.