NVIDIA Skills
by @nvidia · plugin · 211 skills
NVIDIA Skills from NVIDIA/skills.
Install the whole plugin (CLI)
npx skillmds add nvidia/cufolio
npx skillmds add nvidia/hsb-app
npx skillmds add nvidia/hsb-test
npx skillmds add nvidia/rag-eval
npx skillmds add nvidia/rag-perf
npx skillmds add nvidia/hsb-flash
npx skillmds add nvidia/hsb-setup
npx skillmds add nvidia/aiq-deploy
npx skillmds add nvidia/cudaq-guide
npx skillmds add nvidia/aiq-research
npx skillmds add nvidia/nemo-rl-docs
npx skillmds add nvidia/cuopt-install
npx skillmds add nvidia/data-designer
npx skillmds add nvidia/mcore-testing
npx skillmds add nvidia/nv-reason-cxr
npx skillmds add nvidia/nv-segment-ct
npx skillmds add nvidia/rag-blueprint
npx skillmds add nvidia/vss-ask-video
npx skillmds add nvidia/deepstream-dev
npx skillmds add nvidia/holoscan-setup
npx skillmds add nvidia/jetson-package
npx skillmds add nvidia/launch-nemo-rl
npx skillmds add nvidia/mcore-split-pr
npx skillmds add nvidia/nemo-retriever
npx skillmds add nvidia/nv-generate-mr
npx skillmds add nvidia/tao-run-automl
npx skillmds add nvidia/tao-train-dino
npx skillmds add nvidia/tao-train-reid
npx skillmds add nvidia/cuopt-developer
npx skillmds add nvidia/nemotron-speech
npx skillmds add nvidia/nv-segment-ctmr
npx skillmds add nvidia/tao-run-on-brev
npx skillmds add nvidia/cuopt-user-rules
npx skillmds add nvidia/cupynumeric-hdf5
npx skillmds add nvidia/jetson-link-docs
npx skillmds add nvidia/jetson-llm-serve
npx skillmds add nvidia/tao-run-deft-aoi
npx skillmds add nvidia/tao-run-on-slurm
npx skillmds add nvidia/tao-run-platform
npx skillmds add nvidia/tao-train-ocdnet
npx skillmds add nvidia/tao-train-ocrnet
npx skillmds add nvidia/tao-train-rtdetr
npx skillmds add nvidia/dali-dynamic-mode
npx skillmds add nvidia/jetson-diagnostic
npx skillmds add nvidia/jetson-init-image
npx skillmds add nvidia/jetson-set-target
npx skillmds add nvidia/tao-finetune-clip
npx skillmds add nvidia/tao-run-on-lepton
npx skillmds add nvidia/vss-manage-alerts
npx skillmds add nvidia/jetson-flash-image
npx skillmds add nvidia/jetson-init-source
npx skillmds add nvidia/jetson-generate-kb
npx skillmds add nvidia/jetson-init-target
npx skillmds add nvidia/jetson-quick-start
npx skillmds add nvidia/mcore-create-issue
npx skillmds add nvidia/mcore-run-on-slurm
npx skillmds add nvidia/nemotron-customize
npx skillmds add nvidia/tao-train-nvdinov2
npx skillmds add nvidia/tao-train-sparse4d
npx skillmds add nvidia/vss-deploy-profile
npx skillmds add nvidia/vss-search-archive
npx skillmds add nvidia/cupynumeric-install
npx skillmds add nvidia/dynamo-troubleshoot
npx skillmds add nvidia/cuopt-server-common
npx skillmds add nvidia/jetson-memory-audit
npx skillmds add nvidia/nemoclaw-user-guide
npx skillmds add nvidia/tao-launch-workflow
npx skillmds add nvidia/jetson-build-source
npx skillmds add nvidia/tao-mine-aoi-images
npx skillmds add nvidia/jetson-download-bsp
npx skillmds add nvidia/tao-train-bevfusion
npx skillmds add nvidia/tao-train-oneformer
npx skillmds add nvidia/tao-train-segformer
npx skillmds add nvidia/vss-query-analytics
npx skillmds add nvidia/vss-summarize-video
npx skillmds add nvidia/dynamo-recipe-runner
npx skillmds add nvidia/earth2studio-install
npx skillmds add nvidia/jetson-customize-fan
npx skillmds add nvidia/jetson-customize-usb
npx skillmds add nvidia/jetson-headless-mode
npx skillmds add nvidia/jetson-llm-benchmark
npx skillmds add nvidia/jetson-promote-image
npx skillmds add nvidia/nv-generate-ct-rflow
npx skillmds add nvidia/nv-generate-mr-brain
npx skillmds add nvidia/physicsnemo-discover
npx skillmds add nvidia/skill-card-generator
npx skillmds add nvidia/tao-train-centerpose
npx skillmds add nvidia/dynamo-router-starter
npx skillmds add nvidia/jetson-customize-mgbe
npx skillmds add nvidia/jetson-customize-pcie
npx skillmds add nvidia/jetson-customize-uphy
npx skillmds add nvidia/jetson-print-bsp-info
npx skillmds add nvidia/cuopt-skill-evolution
npx skillmds add nvidia/jetson-validate-image
npx skillmds add nvidia/nemo-evaluator-plugin
npx skillmds add nvidia/earth2studio-discover
npx skillmds add nvidia/nemo-rl-auto-research
npx skillmds add nvidia/tao-list-capabilities
npx skillmds add nvidia/tao-run-on-kubernetes
npx skillmds add nvidia/tao-train-mask2former
npx skillmds add nvidia/jetson-derive-carrier
npx skillmds add nvidia/tao-train-single-step
npx skillmds add nvidia/tilegym-cutile-python
npx skillmds add nvidia/dicom-metadata-extract
npx skillmds add nvidia/dicom-series-preflight
npx skillmds add nvidia/dicom-series-to-volume
npx skillmds add nvidia/holoscan-install-conda
npx skillmds add nvidia/holoscan-install-wheel
npx skillmds add nvidia/jetson-optimize-memory
npx skillmds add nvidia/nemo-rl-brev-etiquette
npx skillmds add nvidia/nemo-rl-session-memory
npx skillmds add nvidia/nv-segment-ct-finetune
npx skillmds add nvidia/tao-train-nvpanoptix3d
npx skillmds add nvidia/tao-train-pointpillars
npx skillmds add nvidia/cuopt-server-api-python
npx skillmds add nvidia/holoscan-install-debian
npx skillmds add nvidia/holoscan-install-source
npx skillmds add nvidia/jetson-customize-camera
npx skillmds add nvidia/jetson-customize-clocks
npx skillmds add nvidia/jetson-customize-pinmux
npx skillmds add nvidia/nemo-mbridge-resiliency
npx skillmds add nvidia/tao-run-on-local-docker
npx skillmds add nvidia/cuopt-routing-api-python
npx skillmds add nvidia/earth2studio-data-fetch
npx skillmds add nvidia/jetson-print-device-info
npx skillmds add nvidia/nv-generate-vae-finetune
npx skillmds add nvidia/tao-analyze-gaps-vlm-bcq
npx skillmds add nvidia/tao-train-grounding-dino
npx skillmds add nvidia/amc-run-video-calibration
npx skillmds add nvidia/dynamo-interconnect-check
npx skillmds add nvidia/jetson-customize-nvpmodel
npx skillmds add nvidia/jetson-inference-mem-tune
npx skillmds add nvidia/nemo-data-designer-plugin
npx skillmds add nvidia/nemotron-policy-generator
npx skillmds add nvidia/omniverse-cad-to-simready
npx skillmds add nvidia/omniverse-realtime-viewer
npx skillmds add nvidia/tao-analyze-changenet-rca
npx skillmds add nvidia/cuopt-routing-formulation
npx skillmds add nvidia/tao-finetune-cosmos-embed
npx skillmds add nvidia/tao-run-inference-service
npx skillmds add nvidia/tao-setup-nvidia-gpu-host
npx skillmds add nvidia/tao-train-deformable-detr
npx skillmds add nvidia/tao-train-mask-auto-label
npx skillmds add nvidia/tilegym-cutile-autotuning
npx skillmds add nvidia/vss-generate-video-report
npx skillmds add nvidia/accelerated-computing-cudf
npx skillmds add nvidia/amc-run-sample-calibration
npx skillmds add nvidia/holoscan-install-container
npx skillmds add nvidia/nemotron-retrieval-recipes
npx skillmds add nvidia/tao-convert-dataset-format
npx skillmds add nvidia/tao-finetune-cosmos-reason
npx skillmds add nvidia/tao-train-visual-changenet
npx skillmds add nvidia/vss-deploy-video-embedding
npx skillmds add nvidia/amc-setup-calibration-stack
npx skillmds add nvidia/deepstream-profile-pipeline
npx skillmds add nvidia/jetson-speculative-decoding
npx skillmds add nvidia/tao-train-depth-anything-v2
npx skillmds add nvidia/tao-train-foundation-stereo
npx skillmds add nvidia/tao-train-mask-auto-encoder
npx skillmds add nvidia/tao-port-huggingface-model
npx skillmds add nvidia/vss-deploy-dense-captioning
npx skillmds add nvidia/vss-manage-video-io-storage
npx skillmds add nvidia/deepstream-generate-pipeline
npx skillmds add nvidia/mcore-linting-and-formatting
npx skillmds add nvidia/tao-generate-image-grounding
npx skillmds add nvidia/tao-run-automl-deft-pipeline
npx skillmds add nvidia/tao-train-action-recognition
npx skillmds add nvidia/tao-validate-dataset-format
npx skillmds add nvidia/tao-train-optical-inspection
npx skillmds add nvidia/tilegym-adding-cutile-kernel
npx skillmds add nvidia/vss-setup-behavior-analytics
npx skillmds add nvidia/nemo-mbridge-multi-node-slurm
npx skillmds add nvidia/nemo-mbridge-perf-cuda-graphs
npx skillmds add nvidia/nv-generate-mr-brain-finetune
npx skillmds add nvidia/tao-train-mask-grounding-dino
npx skillmds add nvidia/tao-train-pose-classification
npx skillmds add nvidia/vss-setup-video-analytics-api
npx skillmds add nvidia/cupynumeric-parallel-data-load
npx skillmds add nvidia/deepstream-import-vision-model
npx skillmds add nvidia/earth2studio-create-datasource
npx skillmds add nvidia/earth2studio-create-diagnostic
npx skillmds add nvidia/earth2studio-create-prognostic
npx skillmds add nvidia/nemo-automodel-launcher-config
npx skillmds add nvidia/tao-finetune-huggingface-model
npx skillmds add nvidia/tao-train-image-classification
npx skillmds add nvidia/vss-generate-video-calibration
npx skillmds add nvidia/cupynumeric-migration-readiness
npx skillmds add nvidia/nemo-automodel-model-onboarding
npx skillmds add nvidia/nemo-mbridge-perf-megatron-fsdp
npx skillmds add nvidia/nemo-mbridge-perf-memory-tuning
npx skillmds add nvidia/nemo-mbridge-recipe-recommender
npx skillmds add nvidia/cuopt-numerical-optimization-api
npx skillmds add nvidia/digital-health-clinical-asr-eval
npx skillmds add nvidia/nemo-mbridge-mlm-bridge-training
npx skillmds add nvidia/nemo-mbridge-perf-cpu-offloading
npx skillmds add nvidia/omniverse-usd-performance-tuning
npx skillmds add nvidia/tao-train-fast-foundation-stereo
npx skillmds add nvidia/vss-deploy-detection-tracking-2d
npx skillmds add nvidia/vss-deploy-detection-tracking-3d
npx skillmds add nvidia/cuopt-multi-objective-exploration
npx skillmds add nvidia/digital-health-clinical-asr-build
npx skillmds add nvidia/digital-health-clinical-asr-setup
npx skillmds add nvidia/nemo-automodel-recipe-development
npx skillmds add nvidia/physical-ai-neural-reconstruction
npx skillmds add nvidia/tao-analyze-gaps-visual-changenet
npx skillmds add nvidia/cuopt-numerical-optimization-api-c
npx skillmds add nvidia/nemo-mbridge-perf-moe-comm-overlap
npx skillmds add nvidia/nemo-mbridge-perf-moe-long-context
npx skillmds add nvidia/nemo-mbridge-perf-moe-vlm-training
npx skillmds add nvidia/nemo-mbridge-perf-sequence-packing
npx skillmds add nvidia/deepstream-sopSkills in this plugin
- ▌ cufolio · nvidia bundleBuild, optimize, backtest, rebalance, or analyze stock portfolios using NVIDIA-accelerated Mean-CVaR optimization with cuOpt GPU solver.
- ▌ hsb-app · nvidia bundleDiscover, select, and run Holoscan Sensor Bridge example applications on a connected devkit over SSH, filtering by platform, board type, and sensors.
- ▌ hsb-test · nvidia bundleExecute QA test plans on Holoscan Sensor Bridge hardware by reading a test document, filtering tests by setup, running automatable tests with pass/fail evaluation, and producing a structured report.
- ▌ rag-eval · nvidia bundleEvaluates RAG pipelines using a filesystem-based benchmark with corpus/ and train.json, running evaluate_rag.py to tune retrieval and generation flags and interpret RAGAS metrics.
- ▌ rag-perf · nvidia bundleRun config-driven performance benchmarks against a deployed NVIDIA RAG Blueprint server, including profiling and load testing, with a unified report.
- ▌ hsb-flash · nvidia bundleFlash FPGA firmware on HSB Lattice boards and Leopard Imaging VB1940 cameras connected to NVIDIA devkits, with safety checks and multi-step upgrade/downgrade procedures.
- ▌ hsb-setup · nvidia bundleSet up the Holoscan Sensor Bridge demo environment end-to-end: clone the repo, configure the host per platform, build and run the demo container, and verify connectivity to the sensor board.
- ▌ aiq-deploy · nvidia bundleInstalls, deploys, runs, validates, troubleshoots, and stops NVIDIA AI-Q Blueprint infrastructure for local or self-hosted servers.
- ▌ cudaq-guide · nvidia bundleGuide users through installing CUDA-Q, writing quantum kernels, running GPU-accelerated simulations, connecting to QPU hardware, and exploring built-in applications.
- ▌ aiq-research · nvidia bundleRuns deep research queries through a reachable NVIDIA AI-Q Blueprint backend, handling health checks, job submission, polling, and report retrieval.
- ▌ nemo-rl-docs · nvidia bundleUpdate docs/index.md and write Google-style docstrings for NeMo-RL documentation changes.
- ▌ cuopt-install · nvidia bundleInstall cuOpt for Python, C, or REST server via pip, conda, or Docker, and verify the installation.
- ▌ data-designer · nvidia bundleBuild synthetic datasets and data generation pipelines using the Data Designer library.
- ▌ mcore-testing · nvidia bundleGuides testing Megatron-LM: test layout, recipe YAML, adding and running unit/functional tests, golden values, marker filters, and CI parity.
- ▌ nv-reason-cxr · nvidia bundleRuns chest X-ray reasoning smoke tests using the NV-Reason-CXR-3B model via local inference or a public Hugging Face Space API.
- ▌ nv-segment-ct · nvidia bundleSegments abdominal organs from CT NIfTI volumes using the NV-Segment-CT VISTA3D model, producing label maps and structured evidence JSON.
- ▌ rag-blueprint · nvidia bundleDeploy, configure, troubleshoot, and manage NVIDIA RAG Blueprint deployments across Docker, Helm, and library setups.
- ▌ vss-ask-video · nvidia bundleAsk visual questions about video clips using a VSS agent's video_understanding tool, requiring a fresh look at frames rather than prior metadata or search results.
- ▌ deepstream-dev · nvidia bundleBuild video analytics pipelines using NVIDIA DeepStream SDK 9.0 with Python pyservicemaker API, including GStreamer-based video processing, TensorRT inference integration, object detection/tracking, and Kafka/message broker integration.
- ▌ holoscan-setup · nvidia bundleInspects the host system, assesses platform compatibility, and recommends the correct Holoscan SDK installation method, then delegates to a method-specific install skill.
- ▌ jetson-package · nvidia bundleSelects Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes based on Orin SM 8.7 vs Thor SM 11.0 and JetPack version.
- ▌ launch-nemo-rl · nvidia bundleLaunch, monitor, stop, and debug NeMo-RL recipes on a Kubernetes cluster using the nrl-k8s CLI, supporting ephemeral and long-lived RayCluster modes.
- ▌ mcore-split-pr · nvidia bundleSplit a large pull request into multiple smaller PRs to reduce the number of required CODEOWNERS reviewer groups.
- ▌ nemo-retriever · nvidia bundleIndex folders of PDFs and other documents into LanceDB for vector search, then query them with semantic search, page filters, verbatim quotes, and cross-document aggregation.
- ▌ nv-generate-mr · nvidia bundleGenerates synthetic body MRI volumes using NVIDIA's NV-Generate-CTMR rflow-mr model. Wraps the upstream diffusion inference pipeline with config staging, output validation, and NIfTI volume summarization.
- ▌ tao-run-automl · nvidia bundleRun automated hyperparameter optimization for NVIDIA TAO models using AutoMLRunner, supporting multiple search algorithms and experiment tracking.
- ▌ tao-train-dino · nvidia bundleTrain, evaluate, export, distill, quantize, or run inference for a TAO DINO 2D object detector using transformer-based detection with denoising training and multi-scale features.
- ▌ tao-train-reid · nvidia bundleTrains, evaluates, exports, and runs inference for person re-identification models using TAO, learning discriminative embeddings for cross-camera matching.
- ▌ cuopt-developer · nvidia bundleModify, build, test, debug, and contribute to the NVIDIA cuOpt codebase (C++/CUDA, Python, server, CI). Includes guidance for solver internals, pull requests, DCO signoff, and code conventions.
- ▌ nemotron-speech · nvidia bundleRoutes NVIDIA Nemotron Speech (Riva) NIM tasks for ASR, TTS, and NMT, covering cloud-hosted inference, self-hosted Docker deployment, and custom model builds.
- ▌ nv-segment-ctmr · nvidia bundleRuns NV-Segment-CTMR segmentation on CT or MRI NIfTI volumes and records label-map evidence.
- ▌ tao-run-on-brev · nvidia bundleManage NVIDIA Brev GPU instances for TAO training, evaluation, and inference using the Brev CLI and Docker.
- ▌ cuopt-user-rules · nvidia bundleGuides users through installing, configuring, and calling the NVIDIA cuOpt SDK for routing and optimization problems, with emphasis on clarifying requirements and safe execution.
- ▌ cupynumeric-hdf5 · nvidia bundleRead and write large cuPyNumeric arrays to HDF5 files using Legate's parallel, distributed HDF5 I/O.
- ▌ jetson-link-docs · nvidia bundleBind pre-downloaded Jetson reference docs (developer guide, design guide, pinmux, schematics) into the active profile documents block. Use after staging docs on disk; not for downloading.
- ▌ jetson-llm-serve · nvidia bundleServe LLMs and VLMs on NVIDIA Jetson devices using vLLM or SGLang with optimized Docker containers and quantization presets.
- ▌ tao-run-deft-aoi · nvidia bundleAutomates the full DEFT AOI improvement loop for NVIDIA TAO VisualChangeNet / ChangeNet PCB inspection models, including baseline evaluation, RCA, synthetic defect generation, data mining, retraining, and deployment gating until KPI targets are met.
- ▌ tao-run-on-slurm · nvidia bundleSubmit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage.
- ▌ tao-run-platform · nvidia bundleSubmit and monitor GPU training jobs on Brev, SLURM, Docker, or Kubernetes using the TAO Execution SDK, with job handles, S3 I/O wrapping, and multi-node distributed training.
- ▌ tao-train-ocdnet · nvidia bundleTrains, evaluates, exports, prunes, quantizes, retrains, and runs inference for OCDNet scene text detection models using TAO, detecting arbitrary-oriented text regions in natural images.
- ▌ tao-train-ocrnet · nvidia bundleTrains, evaluates, exports, prunes, quantizes, retrains, and runs inference for TAO OCRNet models for scene text recognition from cropped text-region images, supporting CTC and attention-based decoders.
- ▌ tao-train-rtdetr · nvidia bundleTrain, evaluate, distill, quantize, export, and run inference for RT-DETR object detection models using NVIDIA TAO.
- ▌ dali-dynamic-mode · nvidia bundleWrite, review, and migrate code using NVIDIA DALI's imperative dynamic-mode API for efficient data loading and preprocessing.
- ▌ jetson-diagnostic · nvidia bundleCaptures a read-only health snapshot from a Jetson device, reporting identity, memory, GPU, thermal, power, storage, services, and top processes.
- ▌ jetson-init-image · nvidia bundleExtract Jetson Linux BSP and sample-rootfs tarballs, run apply_binaries.sh with the correct GPU stack flag, and record the image metadata in the active target profile.
- ▌ jetson-set-target · nvidia bundleSwitch the active Jetson target-platform pointer to an existing profile YAML. Use before customize/build/flash to change target; not for authoring profiles — use jetson-init-target instead.
- ▌ tao-finetune-clip · nvidia bundleFine-tune and deploy CLIP vision-language models for zero-shot classification, image-text retrieval, and embedding extraction with ONNX and TensorRT support.
- ▌ tao-run-on-lepton · nvidia bundleSubmit TAO jobs to Lepton managed GPU compute on DGX Cloud, with run/status/cancel interface and multi-node distributed training support.
- ▌ vss-manage-alerts · nvidia bundleOperate the VSS alert pipeline for real-time monitoring, Alert-Bridge subscriptions, Slack notifications, incident queries, and camera onboarding.
- ▌ jetson-flash-image · nvidia bundleFlash a promoted BSP image to a Jetson device in RCM mode using NVIDIA's flash.sh or l4t_initrd_flash.sh toolchain.
- ▌ jetson-init-source · nvidia bundleBootstraps a BSP customization workspace for NVIDIA Jetson, including the Linux_for_Tegra overlay tracker, bsp_sources mono-tree, and Crosstool-NG toolchain.
- ▌ jetson-generate-kb · nvidia bundleGenerates a per-target knowledge-base markdown file by walking the BSP root and source tree, documenting image layout, source tree structure, and document references.
- ▌ jetson-init-target · nvidia bundleCreates a new Jetson target-platform profile YAML by selecting a reference devkit and optional custom carrier, then updates the active target pointer.
- ▌ jetson-quick-start · nvidia bundleDispatches Jetson BSP customization by presenting a click-to-select setup questionnaire and passing prefilled answers to downstream setup skills.
- ▌ mcore-create-issue · nvidia bundleInvestigate a failing GitHub Actions job, extract the root cause, and file a well-structured bug issue against NVIDIA/Megatron-LM.
- ▌ mcore-run-on-slurm · nvidia bundleLaunch distributed Megatron-LM training jobs on a SLURM cluster with a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules, container conventions, monitoring, and per-rank failure diagnosis.
- ▌ nemotron-customize · nvidia bundlePlan, configure, and chain Nemotron model customization steps into single-step or multi-step pipelines for curation, translation, fine-tuning, RL alignment, benchmarking, checkpoint conversion, optimization, and evaluation.
- ▌ tao-train-nvdinov2 · nvidia bundleTrains vision transformers via self-distillation without labels for self-supervised visual representation learning, and supports export and inference of NVDINOv2 backbones.
- ▌ tao-train-sparse4d · nvidia bundleTrains, evaluates, exports, quantizes, and runs inference for Sparse4D multi-camera temporal 3D object detection and tracking models using TAO.
- ▌ vss-deploy-profile · nvidia bundleSelects, configures, deploys, verifies, debugs, or tears down a VSS profile (base, search, lvs, warehouse, edge) for NVIDIA's video search and summarization stack.
- ▌ vss-search-archive · nvidia bundleSearch archived video using natural language, ingest video files or RTSP streams, and manage ingested sources.
- ▌ cupynumeric-install · nvidia bundleInstall and verify cuPyNumeric for Python using conda or pip, including GPU usage checks.
- ▌ dynamo-troubleshoot · nvidia bundleDiagnose failed or unhealthy Dynamo deployments by collecting a read-only debug bundle, classifying failures, and providing step-by-step remediation guidance.
- ▌ cuopt-server-common · nvidia bundleExplains the domain concepts of the cuOpt REST server, including supported problem types, request flow, and key endpoints, without deployment or client code.
- ▌ jetson-memory-audit · nvidia bundleMeasure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
- ▌ nemoclaw-user-guide · nvidia bundleGuides AI coding assistants to the official NemoClaw documentation via MCP server or Markdown files for installation, configuration, operation, and troubleshooting.
- ▌ tao-launch-workflow · nvidia bundleCollects launch inputs and runs preflight checks before executing TAO workflows such as AutoML, training, evaluation, inference, export, TensorRT engine generation, or DEFT jobs on supported platforms.
- ▌ jetson-build-source · nvidia bundleRebuild kernel-side artifacts (DTBs, OOT modules, kernel Image) from source changes under bsp_sources/ and produce a manifest for staging into a BSP image.
- ▌ tao-mine-aoi-images · nvidia bundleEmbeds target and source image parquets, then mines nearest-neighbour source images for augmentation in VCN AOI workflows.
- ▌ jetson-download-bsp · nvidia bundleDownloads NVIDIA Jetson Linux BSP artifacts (BSP tarball, sample rootfs, public_sources, x-tools, guides) for the active target. Used for Auto Setup; does not extract or edit profiles.
- ▌ tao-train-bevfusion · nvidia bundleTrains, evaluates, and runs inference for BEVFusion multi-sensor 3D object detection models that fuse LiDAR and camera data in bird's-eye-view space for autonomous driving.
- ▌ tao-train-oneformer · nvidia bundleTrain, evaluate, export, quantize, and run inference for a TAO OneFormer model that performs panoptic, instance, and semantic segmentation using task-conditioned queries.
- ▌ tao-train-segformer · nvidia bundleTrains, evaluates, exports, quantizes, and runs inference for SegFormer semantic segmentation models using NVIDIA TAO.
- ▌ vss-query-analytics · nvidia bundleQueries video analytics incidents, alerts, metrics, and sensor data from Elasticsearch via the VA-MCP server.
- ▌ vss-summarize-video · nvidia bundleSummarize recorded video clips using the LVS microservice with a VLM fallback, producing a narrative summary with timestamped events.
- ▌ dynamo-recipe-runner · nvidia bundleSelect, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes for model serving with GPU support.
- ▌ earth2studio-install · nvidia bundleGuides installing Earth2Studio via uv or pip, selecting model extras, and configuring environment variables.
- ▌ jetson-customize-fan · nvidia bundleAdd, remove, edit, list, or change the boot default of an nvfancontrol fan profile on a Jetson/Tegra target.
- ▌ jetson-customize-usb · nvidia bundleEnable, disable, or change the role of USB2/USB3 SS ports on Jetson custom carriers by generating kernel-DT overlays that flip lane, port, and host xHCI phys in lockstep.
- ▌ jetson-headless-mode · nvidia bundlePlan and apply safe, reversible headless-mode changes on Jetson devices to reclaim memory from the GUI and non-essential daemons.
- ▌ jetson-llm-benchmark · nvidia bundleBenchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
- ▌ jetson-promote-image · nvidia bundleCopies overlay files and built artifacts into a staged BSP image for NVIDIA Jetson platforms, preparing it for flashing without modifying the workspace.
- ▌ nv-generate-ct-rflow · nvidia bundleGenerates synthetic CT volumes and masks using NVIDIA's rectified-flow pipeline for medical imaging research.
- ▌ nv-generate-mr-brain · nvidia bundleGenerates synthetic brain MRI volumes using NVIDIA's NV-Generate-CTMR workflow, with configurable modality and random seed.
- ▌ physicsnemo-discover · nvidia bundleNavigate the PhysicsNeMo repository by discovering model families, datapipes, and examples through live file search, without writing training code.
- ▌ skill-card-generator · nvidia bundleGenerates or updates a governance skill card for an existing agent skill directory by analyzing source signals, building a grounded JSON context, and rendering a deterministic markdown card.
- ▌ tao-train-centerpose · nvidia bundleTrain, evaluate, export, and run inference for CenterPose models used in 6-DoF object pose estimation with keypoint regression.
- ▌ dynamo-router-starter · nvidia bundleStart or patch Dynamo router modes and run router endpoint smoke checks for round-robin, KV-aware, least-loaded, or device-aware routing.
- ▌ jetson-customize-mgbe · nvidia bundleGenerates kernel-DT overlay fragments to enable 25G/10G/1G MGBE QSFP interfaces on Jetson Thor, verifying pinmux and integrating with the BSP customization workflow.
- ▌ jetson-customize-pcie · nvidia bundleGenerates kernel device-tree overlay fragments to enable or disable individual PCIe controllers and configure lane count and link speed on Jetson Thor/Orin custom carriers.
- ▌ jetson-customize-uphy · nvidia bundleConfigure Jetson UPHY lane allocation on Orin/Thor custom carriers by editing ODMDATA tokens and dispatching per-controller kernel-DT overlay skills.
- ▌ jetson-print-bsp-info · nvidia bundleInspects a Jetson Linux_for_Tegra BSP tree on the host PC and prints a concise summary including L4T version, board configs, and rootfs state.
- ▌ cuopt-skill-evolution · nvidia bundleDetects generalizable learnings from problem-solving interactions and proposes skill updates to improve future performance.
- ▌ jetson-validate-image · nvidia bundleRun static BSP checks and on-target smoke/regression tests on a flashed NVIDIA Jetson device to validate a customized BSP image.
- ▌ nemo-evaluator-plugin · nvidia bundleRun evaluation tasks against a NeMo Platform server using the Evaluator plugin CLI and Python SDK.
- ▌ earth2studio-discover · nvidia bundleFind Earth2Studio models, data sources, and examples for weather/climate use cases by consulting live documentation and verifying compatibility via the lexicon system.
- ▌ nemo-rl-auto-research · nvidia bundleGuides agents through the full lifecycle of NeMo-RL experiments: understanding recipes, launching reproducible runs, analyzing results, and preserving human oversight with git and TSV logs.
- ▌ tao-list-capabilities · nvidia bundleLists TAO Skill Bank capabilities, models, and AutoML support by running packaged scripts.
- ▌ tao-run-on-kubernetes · nvidia bundleSubmits TAO container jobs as single-pod Kubernetes Jobs with NVIDIA GPU scheduling on EKS, GKE, AKS, or on-prem clusters.
- ▌ tao-train-mask2former · nvidia bundleTrain, evaluate, export, quantize, and run inference on Mask2Former models for panoptic, instance, and semantic segmentation using NVIDIA TAO.
- ▌ jetson-derive-carrier · nvidia bundleBootstrap a custom carrier board by forking carrier files and scaffolding a DT overlay from the reference devkit.
- ▌ tao-train-single-step · nvidia bundleFine-tune a TAO model with standard supervised training, evaluation, and export, with AutoML bypass and platform-specific credential intake.
- ▌ tilegym-cutile-python · nvidia bundleWrite high-performance GPU kernels using cuTile's tile-based programming model with validation and optimization, including deep agent orchestration for complex multi-kernel tasks.
- ▌ dicom-metadata-extract · nvidia bundleExtracts selected metadata from a DICOM file and flags standard-tag PHI presence. Not for anonymization or clinical use.
- ▌ dicom-series-preflight · nvidia bundleScans a DICOM series folder to extract header metadata and produce a preflight verdict without decoding pixel data.
- ▌ dicom-series-to-volume · nvidia bundleConverts a single CT DICOM series folder into a Hounsfield Unit NIfTI volume with affine metadata.
- ▌ holoscan-install-conda · nvidia bundleInstall Holoscan SDK v4.3+ via Conda in a CUDA 13 environment, including Python bindings and C++ development headers.
- ▌ holoscan-install-wheel · nvidia bundleInstall the Holoscan SDK Python wheel via pip into a virtual environment and verify with example scripts.
- ▌ jetson-optimize-memory · nvidia bundleReclaim DRAM on NVIDIA Jetson devices by disabling unused display, camera, and DMA subsystems across MB1 BCT, MB2 BCT, kernel reserved-memory, and SWIOTLB layers for headless or no-camera deployments.
- ▌ nemo-rl-brev-etiquette · nvidia bundleProvides storage and environment conventions for NeMo-RL agents on Brev instances, ensuring large experiment outputs go to /ephemeral and secrets are loaded from .env.
- ▌ nemo-rl-session-memory · nvidia bundleMaintain durable session memory across agent disconnects by writing structured checkpoints to the repo's session directory, enabling context recovery.
- ▌ nv-segment-ct-finetune · nvidia bundleFine-tune NV-Segment-CT VISTA3D on CT NIfTI labels for smoke testing or dataset adaptation, wrapping the upstream MONAI bundle entrypoint.
- ▌ tao-train-nvpanoptix3d · nvidia bundleTrains, evaluates, exports, and runs inference for NVPanoptix3D models that perform panoptic 3D scene reconstruction from posed RGB images, producing 3D panoptic segmentation with occupancy completion.
- ▌ tao-train-pointpillars · nvidia bundleTrain, evaluate, export, prune, and run inference for PointPillars 3D object detection models from LiDAR point clouds using NVIDIA TAO.
- ▌ cuopt-server-api-python · nvidia bundleDeploy a cuOpt REST server and solve routing optimization problems using Python or curl.
- ▌ holoscan-install-debian · nvidia bundleInstall the Holoscan SDK C++ runtime and headers on Ubuntu using NVIDIA's apt repository, with automatic CUDA variant detection and verification via bundled examples.
- ▌ holoscan-install-source · nvidia bundleBuild the Holoscan SDK from source using its in-tree Docker-based build script, producing a local install tree for CMake-based applications.
- ▌ jetson-customize-camera · nvidia bundleEnable MIPI/GMSL camera sensors on a Jetson Thor or Orin custom carrier by rendering a kernel-DT overlay from the in-tree sensor DTSI.
- ▌ jetson-customize-clocks · nvidia bundleLock, cap, or customize CPU, GPU, and EMC clock behavior on NVIDIA Jetson devices by editing BPMP DTB and nvpower.sh before flashing.
- ▌ jetson-customize-pinmux · nvidia bundleParse a Jetson Orin/Thor pinmux XLSM spreadsheet and generate per-pin BCT DTSI files for a custom carrier board.
- ▌ nemo-mbridge-resiliency · nvidia bundleConfigure fault tolerance, straggler detection, preemption, in-process restart, and re-run state machine for Megatron Bridge training jobs.
- ▌ tao-run-on-local-docker · nvidia bundleRun TAO SDK jobs as Docker containers on a local or remote Docker daemon with NVIDIA GPU support, including preflight checks and credential handling.
- ▌ cuopt-routing-api-python · nvidia bundleSolve vehicle routing problems (TSP, VRP, PDP) using NVIDIA cuOpt's Python API with cost matrices, time windows, capacity constraints, and pickup-delivery pairs.
- ▌ earth2studio-data-fetch · nvidia bundleGuides users through downloading weather/climate data via Earth2Studio data source APIs, verifying variable support via the lexicon system, and generating a working Python fetch script.
- ▌ jetson-print-device-info · nvidia bundleCaptures a baseline snapshot of a Jetson device's module model, L4T version, kernel, OS version, and power mode for performance testing or verification.
- ▌ nv-generate-vae-finetune · nvidia bundleFinetune the NV-Generate-CTMR MAISI VAE/autoencoder on user-supplied CT or MRI NIfTI volumes using a staged config and datalist workflow.
- ▌ tao-analyze-gaps-vlm-bcq · nvidia bundleExtract false-positive and false-negative gaps from VLM binary-classification-question predictions by comparing model responses against ground truth, producing a structured JSONL file and summary report for downstream root-cause analysis.
- ▌ tao-train-grounding-dino · nvidia bundleTrains, evaluates, exports, quantizes, and runs inference for a Grounding DINO model that detects objects described by text prompts without a fixed class vocabulary.
- ▌ amc-run-video-calibration · nvidia bundleCalibrate a new dataset from pre-recorded video files via the AutoMagicCalib REST API.
- ▌ dynamo-interconnect-check · nvidia bundleValidates that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after deploying a disagg or multi-node recipe to confirm KV transport is correct, or use troubleshoot for already-failed pods.
- ▌ jetson-customize-nvpmodel · nvidia bundleAdd, remove, edit, list, or change the boot default of nvpmodel power modes on Jetson/Tegra platforms (Orin, Thor) by modifying the BSP-side configuration file.
- ▌ jetson-inference-mem-tune · nvidia bundleRecommends an inference runtime and memory-related launch flags for LLM/VLM workloads on NVIDIA Jetson devices, based on a live memory audit snapshot.
- ▌ nemo-data-designer-plugin · nvidia bundleBuild synthetic datasets and data generation pipelines using the Data Designer library.
- ▌ nemotron-policy-generator · nvidia bundleGenerates custom safety policies for NVIDIA Nemotron content-safety guardrails, producing a Markdown policy, JSON taxonomy, and inference prompts from rough user input.
- ▌ omniverse-cad-to-simready · nvidia bundleCoordinate the end-to-end pipeline from CAD or source assets to SimReady USD assets, including conversion, material and physics assignment, validation, conformance, and optional packaging.
- ▌ omniverse-realtime-viewer · nvidia bundleRoutes Omniverse Realtime Viewer requests to focused reference documents and enforces architectural rules for building USD viewer applications.
- ▌ tao-analyze-changenet-rca · nvidia bundlePerforms deep root cause analysis on NVIDIA TAO Visual ChangeNet classification experiments, using image-evidence-driven investigation to diagnose model failures and produce actionable reports.
- ▌ cuopt-routing-formulation · nvidia bundleDefines vehicle routing problem types (TSP, VRP, PDP) and the data requirements needed to formulate them, without covering any API or interface details.
- ▌ tao-finetune-cosmos-embed · nvidia bundleFine-tune, evaluate, run inference, and export Cosmos-Embed1 video-text embedding models for tasks like text-to-video retrieval and semantic deduplication.
- ▌ tao-run-inference-service · nvidia bundleStart, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
- ▌ tao-setup-nvidia-gpu-host · nvidia bundleChecks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
- ▌ tao-train-deformable-detr · nvidia bundleTrain, evaluate, export, quantize, and run inference for a Deformable DETR 2D object detection model using TAO, with deformable attention for efficient multi-scale feature processing.
- ▌ tao-train-mask-auto-label · nvidia bundleTrains, evaluates, and runs inference for Mask Auto-Label (MAL) weakly-supervised segmentation models using ViT-MAE backbones with minimal point or box annotations.
- ▌ tilegym-cutile-autotuning · nvidia bundleAdds autotuning to CuTile kernels using the exhaustive_search API with a tune-once/cache/direct-launch pattern, covering occupancy-only and complex tile-size search spaces.
- ▌ vss-generate-video-report · nvidia bundleGenerates video analysis reports by routing to a VLM backend for per-clip analysis or an analytics backend for incident-range reports, with deployment profile verification and URL rewriting.
- ▌ accelerated-computing-cudf · nvidia bundleAccelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
- ▌ amc-run-sample-calibration · nvidia bundleRun end-to-end calibration on the bundled sample dataset against a running AMC microservice to verify the stack works before processing real data.
- ▌ holoscan-install-container · nvidia bundlePull and verify the official Holoscan SDK container from NGC, selecting the correct CUDA/arch tag for the host GPU and validating with bundled Python and C++ examples.
- ▌ nemotron-retrieval-recipes · nvidia bundlePlan, debug, tune, evaluate, export, or deploy public Nemotron embedding and reranking retrieval recipes using the current checkout.
- ▌ tao-convert-dataset-format · nvidia bundleConverts NVIDIA TAO DAFT datasets between supported formats using the `tao-daft convert` CLI.
- ▌ tao-finetune-cosmos-reason · nvidia bundleFine-tune Cosmos Reason video QA models using supervised fine-tuning with FSDP parallelism, including dataset preparation, spec construction, and AutoML support.
- ▌ tao-train-visual-changenet · nvidia bundleTrains, evaluates, exports, and runs inference for Visual ChangeNet models used in AOI defect detection, comparing image pairs for PASS/NO_PASS classification or change-segmentation masks.
- ▌ vss-deploy-video-embedding · nvidia bundleDeploy and operate the VSS 3.2 GA RT-Embed Video Embedding microservice using Docker Compose, covering GPU prerequisites, REST API usage for file uploads, text/video embeddings, live RTSP streams, Redis/Kafka/OTel integration, and troubleshooting.
- ▌ amc-setup-calibration-stack · nvidia bundleDeploy the AutoMagicCalib microservice and web UI from pre-built NGC release images using Docker Compose.
- ▌ deepstream-profile-pipeline · nvidia bundleProfile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement.
- ▌ jetson-speculative-decoding · nvidia bundleReduce per-token latency on Jetson vLLM servers by appending speculative decoding configuration, with guidance on when to enable and how to benchmark the improvement.
- ▌ tao-train-depth-anything-v2 · nvidia bundleTrain, evaluate, export, and run inference for monocular depth estimation models using Metric Depth Anything v2 or Relative Depth Anything architectures via the TAO toolkit.
- ▌ tao-train-foundation-stereo · nvidia bundleTrains, evaluates, exports, and runs inference on FoundationStereo models for stereo depth estimation and 3D reconstruction from stereo image pairs.
- ▌ tao-train-mask-auto-encoder · nvidia bundleTrain, evaluate, export, and run inference for Masked Auto-Encoder (MAE) models for self-supervised pretraining and fine-tuning of visual representations.
- ▌ tao-port-huggingface-model · nvidia bundleIntegrate a HuggingFace computer vision model into the NVIDIA TAO Toolkit ecosystem, covering the full pipeline from prerequisites to container testing.
- ▌ vss-deploy-dense-captioning · nvidia bundleDeploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
- ▌ vss-manage-video-io-storage · nvidia bundleManage VIOS and NvStreamer REST API operations for video input/output and storage, including sensors, streams, snapshots, clips, and recordings.
- ▌ deepstream-generate-pipeline · nvidia bundleBuilds and validates DeepStream GStreamer pipelines through an interactive questionnaire and a BM25 retrieval engine over 270+ verified pipelines.
- ▌ mcore-linting-and-formatting · nvidia bundleLint and format Python code for Megatron-LM using ruff, black, isort, pylint, and mypy, with commands for autoformatting and import ordering.
- ▌ tao-generate-image-grounding · nvidia bundleGenerates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
- ▌ tao-run-automl-deft-pipeline · nvidia bundleRuns a three-phase AOI training pipeline: AutoML HPO baseline, DEFT iterative data improvement, and AutoML refinement on the augmented dataset.
- ▌ tao-train-action-recognition · nvidia bundleTrain, evaluate, export, and run inference on TAO action-recognition models for classifying temporal actions in video clips using RGB, optical flow, or joint input.
- ▌ tao-validate-dataset-format · nvidia bundleValidates NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors using the `tao-daft validate` CLI tool.
- ▌ tao-train-optical-inspection · nvidia bundleTrains, evaluates, exports, and runs inference for Siamese-network-based optical inspection models to detect manufacturing defects and quality issues in image pairs.
- ▌ tilegym-adding-cutile-kernel · nvidia bundleAdd a new cuTile GPU kernel operator to TileGym, covering dispatch registration, backend implementation, exports, tests, and benchmarks.
- ▌ vss-setup-behavior-analytics · nvidia bundleDeploy the behavior-analytics service standalone with a chosen entrypoint, config source, and optional calibration, without the full warehouse stack.
- ▌ nemo-mbridge-multi-node-slurm · nvidia bundleConvert single-node PyTorch distributed scripts into multi-node Slurm sbatch jobs and debug common multi-node failures, covering srun-native and torch.distributed approaches, container setup, NCCL timeouts, and interactive allocation.
- ▌ nemo-mbridge-perf-cuda-graphs · nvidia bundleValidate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
- ▌ nv-generate-mr-brain-finetune · nvidia bundleFinetunes the NV-Generate-CTMR MR-brain diffusion UNet from user-supplied NIfTI training volumes using a wrapper that stages configs and delegates to upstream scripts.
- ▌ tao-train-mask-grounding-dino · nvidia bundleTrains, evaluates, exports, quantizes, and runs inference for a Mask Grounding DINO model for open-set instance segmentation guided by text prompts.
- ▌ tao-train-pose-classification · nvidia bundleTrain, evaluate, export, and run inference for pose classification models using ST-GCN on skeleton keypoint sequences.
- ▌ vss-setup-video-analytics-api · nvidia bundleDeploys the vss-video-analytics-api REST service standalone with config, data-log bind, and optional Elasticsearch/Kafka connectivity.
- ▌ cupynumeric-parallel-data-load · nvidia bundleLoad sharded datasets (npy, Parquet, HDF5, raw binary) into distributed cuPyNumeric arrays using manual partitioning and Legate task launches.
- ▌ deepstream-import-vision-model · nvidia bundleImport object detection models from HuggingFace or NVIDIA NGC into a DeepStream pipeline with automated ONNX download, TensorRT engine build, custom parser, multi-stream benchmark, and PDF report generation.
- ▌ earth2studio-create-datasource · nvidia bundleCreate and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores like S3, GCS, Azure, HTTP, or HuggingFace.
- ▌ earth2studio-create-diagnostic · nvidia bundleCreate Earth2Studio diagnostic model wrappers for single-step data transformations, including simple derived diagnostics, packaged AutoModel diagnostics, and generative or diffusion diagnostics.
- ▌ earth2studio-create-prognostic · nvidia bundleCreate Earth2Studio prognostic model wrappers that time-step weather forecasts forward, with triple-inheritance classes, tests, and documentation.
- ▌ nemo-automodel-launcher-config · nvidia bundleConfigure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
- ▌ tao-finetune-huggingface-model · nvidia bundleFine-tune HuggingFace CV, VLM, or LLM models on local NVIDIA GPUs using an NGC PyTorch container, with support for full or LoRA training, dataset handling, and optional model push to the Hub.
- ▌ tao-train-image-classification · nvidia bundleTrain, evaluate, distill, quantize, export, and run inference for PyTorch-based TAO image classification models with support for multiple backbones.
- ▌ vss-generate-video-calibration · nvidia bundleRuns AutoMagicCalib calibration on local MP4s, RTSP streams, or a bundled sample dataset, and deploys the AMC microservice when needed.
- ▌ cupynumeric-migration-readiness · nvidia bundleAssesses NumPy code for cuPyNumeric migration readiness by analyzing source code against an API support manifest and GPU-scaling idioms, producing a structured verdict with per-finding reasoning.
- ▌ nemo-automodel-model-onboarding · nvidia bundleGuides implementation of new model architectures in NeMo AutoModel through five phases: discovery, implementation, registration, validation, and testing.
- ▌ nemo-mbridge-perf-megatron-fsdp · nvidia bundleEnables Megatron Fully Sharded Data Parallel in Megatron-Bridge with configuration overrides, code anchors, pitfalls, and verification steps.
- ▌ nemo-mbridge-perf-memory-tuning · nvidia bundleReduces peak GPU memory in Megatron Bridge training by applying expandable segments, parallelism resizing, activation recompute, and CPU offloading constraints.
- ▌ nemo-mbridge-recipe-recommender · nvidia bundleIndexes Megatron Bridge recipes and recommends the best starting config based on model, GPU count, and training goal.
- ▌ cuopt-numerical-optimization-api · nvidia bundleModel and solve LP, MILP, and QP problems using NVIDIA cuOpt's GPU-accelerated solver via Python, C/C++, or CLI interfaces.
- ▌ digital-health-clinical-asr-eval · nvidia bundleScore a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
- ▌ nemo-mbridge-mlm-bridge-training · nvidia bundleRun Megatron-LM (MLM) and Megatron Bridge training with mock or real data, covering correlation testing, available recipes, and multi-GPU examples.
- ▌ nemo-mbridge-perf-cpu-offloading · nvidia bundleConfigure and validate CPU offloading for Megatron Bridge training, including activation offloading and optimizer state offloading with HybridDeviceOptimizer.
- ▌ omniverse-usd-performance-tuning · nvidia bundleDiagnose and optimize slow-loading, high-memory, or low-FPS USD scenes using a structured workflow with profiling, validation, and mutation phases.
- ▌ tao-train-fast-foundation-stereo · nvidia bundleTrains, evaluates, exports, and runs inference for FastFoundationStereo (FFS) stereo depth estimation models, a distilled variant of FoundationStereo with lower latency.
- ▌ vss-deploy-detection-tracking-2d · nvidia bundleDeploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
- ▌ vss-deploy-detection-tracking-3d · nvidia bundleDeploy and operate the RTVI-CV-3D microservice for multi-camera 3D detection and tracking, supporting sample datasets, custom videos, and RTSP streams.
- ▌ cuopt-multi-objective-exploration · nvidia bundleTrace and interpret the Pareto frontier across competing objectives using repeated single-objective cuOpt solves (weighted-sum and ε-constraint).
- ▌ digital-health-clinical-asr-build · nvidia bundleCurates clinical-specialty term lists, generates IPA-tagged synthetic audio via TTS, and produces NeMo-format manifests for ASR benchmark evaluation.
- ▌ digital-health-clinical-asr-setup · nvidia bundleBootstraps a clinical ASR evaluation environment by verifying NVIDIA_API_KEY, installing Python dependencies, and running a smoke test against hosted TTS/ASR services.
- ▌ nemo-automodel-recipe-development · nvidia bundleCreate and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.
- ▌ physical-ai-neural-reconstruction · nvidia bundleRoutes NuRec/Neural Reconstruction requests to the correct upstream NVIDIA skill (datasets, conversion, training, rendering, object harvesting, frame cleanup).
- ▌ tao-analyze-gaps-visual-changenet · nvidia bundleIdentifies the weakest samples per ground-truth label in NVIDIA TAO VCN Classify experiments by running a Docker container that performs threshold sweep, weakness scoring, and per-lighting expansion, then surfaces top-K weak samples for downstream augmentation or relabeling.
- ▌ cuopt-numerical-optimization-api-c · nvidia bundleSolve LP, MILP, and QP problems using the cuOpt C API with a consistent build pattern and core calls.
- ▌ nemo-mbridge-perf-moe-comm-overlap · nvidia bundleOptimizes MoE expert-parallel communication overlap in Megatron Bridge, covering dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
- ▌ nemo-mbridge-perf-moe-long-context · nvidia bundleProvides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
- ▌ nemo-mbridge-perf-moe-vlm-training · nvidia bundleProvides practical guidance for training Mixture-of-Experts Vision-Language Models in Megatron Bridge, comparing FSDP and 3D-parallel approaches with lessons from recent multimodal experiments.
- ▌ nemo-mbridge-perf-sequence-packing · nvidia bundleValidate and configure packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs with correct context parallelism constraints.
- ▌ deepstream-sop · nvidia bundleBuild, deploy, evaluate, debug, and measure latency for a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection and VLM classification.