all publishers

NVIDIA

@nvidia source repo

962 published skills · page 3 of 10

  1. ▌
    Vss Manage Alerts 2 · nvidia bundle
    Use for VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, Slack notifications, incident queries, camera onboarding. Not for non-alert analytics.
    2.2k repo stars
  2. ▌
    Mcore Create Issue 2 · nvidia bundle
    Investigate a failing GitHub Actions run or job and create a GitHub issue for the failure.
    2.2k repo stars
  3. ▌
    Mcore Run On Slurm 2 · nvidia bundle
    How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis.
    2.2k repo stars
  4. ▌
    Nemotron Customize 2 · nvidia bundle
    Plan, configure, and chain repo-native Nemotron customization steps into single-step or multi-step pipelines: curation, translation, SFT/PEFT (AutoModel or Megatron-Bridge), pretraining/CPT, RL alignment (DPO/RLVR/GRPO/RLHF), BYOB/MCQ benchmarks, checkpoint conversion, ModelOpt optimization, env profiles, and evaluation of trained checkpoints or existing/hosted endpoints. Use when a request names a Nemotron step or workflow, or asks to clean, translate, train, fine-tune, align, convert, optimize, evaluate, or compose these into a pipeline. Do NOT use for frontend/dashboard/visualization work, generic ML advice, billing/access, or non-Nemotron coding tasks.
    2.2k repo stars
  5. ▌
    Vss Search Archive 2 · nvidia bundle
    Use this skill to run top-level VSS fusion search on archived video, or to ingest video files / RTSP streams for search. Do NOT use for ad-hoc visual Q&A (use vss-ask-video), live captioning (use vss-deploy-dense-captioning), or video summarization and reports (use vss-summarize-video).
    2.2k repo stars
  6. ▌
    Cupynumeric Install 2 · nvidia bundle
    Install and verify cuPyNumeric for Python — requirements, commands, verification. Source builds are out of scope.
    2.2k repo stars
  7. ▌
    Dynamo Troubleshoot 2 · nvidia bundle
    Diagnose failed or unhealthy Dynamo deployments. Use when pods, model-cache jobs, PVCs, workers, frontend/router health, endpoints, or benchmark jobs fail; use recipe-runner/router-starter before this for normal bring-up.
    2.2k repo stars
  8. ▌
    Vss Query Analytics 2 · nvidia bundle
    Use this skill when reading video-analytics metrics, incidents, alerts, and sensor data via the VA-MCP server (port 9901). Not for live VLM or incident-range narrative reports.
    2.2k repo stars
  9. ▌
    Dynamo Recipe Runner 2 · nvidia bundle
    Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes. Use for model/backend/GPU/deployment-mode recipe bring-up; use router-starter for router-only mode work and troubleshoot for broken deployments.
    2.2k repo stars
  10. ▌
    Earth2studio Install 2 · nvidia bundle
    Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. Do NOT use for writing inference code, choosing models, or PhysicsNeMo questions.
    2.2k repo stars
  11. ▌
    Nv Generate Ct Rflow 2 · nvidia bundle
    Used for generating synthetic CT volumes and masks with NV-Generate-CTMR rflow-ct. Not for production training data without review.
    2.2k repo stars
  12. ▌
    Nv Generate Mr Brain 2 · nvidia bundle
    Used for generating synthetic T1, T2, FLAIR, SWI, or MRA brain MRI volumes with NV-Generate-CTMR MR-Brain v1. Not for production training data.
    2.2k repo stars
  13. ▌
    Earth2studio Discover 2 · nvidia bundle
    Find Earth2Studio models, data sources, and examples for a weather/climate use case. Do NOT use for writing inference code, downloading data, or installation.
    2.2k repo stars
  14. ▌
    Dicom Metadata Extract 2 · nvidia bundle
    Used for extracting selected metadata from one DICOM file and flagging standard-tag PHI presence. Not for anonymization or clinical use.
    2.2k repo stars
  15. ▌
    Dicom Series Preflight 2 · nvidia bundle
    Used for header-only preflight of one DICOM series folder before conversion or inference. Not for de-identification or clinical clearance.
    2.2k repo stars
  16. ▌
    Dicom Series To Volume 2 · nvidia bundle
    Used for converting one CT DICOM series folder to a HU NIfTI volume with affine evidence. Not for multi-frame DICOM or clinical use.
    2.2k repo stars
  17. ▌
    Holoscan Install Conda 2 · nvidia bundle
    Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment. Use for Conda installs; redirect CUDA 12 hosts to container/wheel.
    2.2k repo stars
  18. ▌
    Holoscan Install Wheel 2 · nvidia bundle
    Install Holoscan SDK Python wheel via pip into a venv. Use for Python installs; not for native C++/apt or Conda installs.
    2.2k repo stars
  19. ▌
    Nv Segment Ct Finetune 2 · nvidia bundle
    Runs standard or fixed-channel softmax finetuning of NV-Segment-CT VISTA3D on CT NIfTI image/label datasets, with optional MONAI-native MLflow tracking and checkpoint evidence. Uses softmax for predefined, mutually exclusive classes; keeps the standard workflow when point prompts or runtime-variable classes are needed. Not for clinical validation.
    2.2k repo stars
  20. ▌
    Cuopt Server API Python 2 · nvidia bundle
    cuOpt REST server — start server, endpoints, Python/curl client examples. Use when the user is deploying or calling the REST API.
    2.2k repo stars
  21. ▌
    Earth2studio Data Fetch 2 · nvidia bundle
    Fetch weather/climate data via Earth2Studio data sources for specific variables and times. Do NOT use for inference pipelines, model discovery, or installation.
    2.2k repo stars
  22. ▌
    Holoscan Install Debian 2 · nvidia bundle
    Install Holoscan SDK natively on Ubuntu via apt. Use for C++ installs on Ubuntu; pair with /holoscan-install-wheel for Python.
    2.2k repo stars
  23. ▌
    Holoscan Install Source 2 · nvidia bundle
    Build Holoscan SDK from source via the in-tree ./run script. Use only when published packages don't meet the user's needs.
    2.2k repo stars
  24. ▌
    Cuopt Routing API Python 2 · nvidia bundle
    Vehicle routing (VRP, TSP, PDP) with cuOpt — Python API only. Use when the user is building or solving routing in Python.
    2.2k repo stars
  25. ▌
    Nv Generate Vae Finetune 2 · nvidia bundle
    Used for finetuning the NV-Generate-CTMR MAISI VAE from CT/MRI NIfTI datalists. Not for clinical or production data approval.
    2.2k repo stars
  26. ▌
    Cicd · nvidia
    CI/CD reference for NeMo-RL. Covers GitHub Actions pipeline structure, CI triggering via /ok to test, and CI failure investigation.
    2.2k repo stars
  27. ▌
    Copyright · nvidia
    NVIDIA copyright header requirements for NeMo-RL. Covers which files need headers and the exact header text.
    2.2k repo stars
  28. ▌
    Review Pr · nvidia bundle
    Interactive PR Review — NVIDIA-NeMo/RL
    2.2k repo stars
  29. ▌
    Contributing · nvidia
    Contribution conventions for NeMo-RL. Covers PR title format, commit sign-off, and CI triggering.
    2.2k repo stars
  30. ▌
    Error Handling · nvidia
    Error handling guidelines for NeMo-RL. Covers exception specificity, minimal try bodies, and else blocks.
    2.2k repo stars
  31. ▌
    Review Pr Team · nvidia
    Agent-team-based parallel code review for NVIDIA-NeMo/RL pull requests. Spawns specialized agents (RL expert, submodule experts, bug finder, design reviewer, test agent, devil's advocate, comment reviewer) that coordinate via shared task list and direct messaging. Leader orchestrates, collates ALL findings, and presents to user for approval before posting.
    2.2k repo stars
  32. ▌
    Config Conventions · nvidia
    Configuration conventions for NeMo-RL. YAML is the single source of truth for defaults. Covers BaseModel/TypedDict usage, dataclass for internal classes, exemplar YAML updates, and forbidden default patterns.
    2.2k repo stars
  33. ▌
    Build And Dependency · nvidia
    Build and dependency management for NeMo-RL. Covers Docker image building and running, uv usage, venv setup, and adding dependencies.
    2.2k repo stars
  34. ▌
    Linting And Formatting · nvidia
    Code style guidelines for NeMo-RL (Python and shell). Covers naming, indentation, comments, docstrings, reflection avoidance, and uv usage.
    2.2k repo stars
  35. ▌
    Rebot Orin Setup 2 · nvidia
    Use when initializing a reBot B601-RS (RobStride 6-DOF, seven motors) arm from an NVIDIA Jetson AGX Orin using the Orin's built-in CAN (mttcan) via an external SN65HVD230 transceiver — no USB-CAN/PCAN adapter. Covers the uv/motorbridge env, bringing up can0 @ 1 Mbit, verifying the bus, reboot persistence, motor-ID assignment, zero calibration, and the CAN bring-up troubleshooting tree (loopback-passes-but-no-motors, pinmux, motor power).
    2.2k repo stars
  36. ▌
    Cosmos3 Codebase Nav 2 · nvidia
    Navigate the Cosmos3 package codebase to find where parameters, configs, defaults, scripts, and documentation live. Use when the user asks "where is X in cosmos3", "how do I find the config for Y", "where are the defaults", "where do I change a parameter", or any question about locating files, modules, or settings. Also use when the user opens or edits files and needs orientation.
    2.2k repo stars
  37. ▌
    Cosmos3 Env Troubleshoot 2 · nvidia
    Diagnose and fix Cosmos3 environment, installation, and runtime errors. Use when the user encounters an ImportError, ModuleNotFoundError, CUDA error, Docker error, checkpoint download failure, or any traceback during setup or inference.
    2.2k repo stars
  38. ▌
    Retail API 2 · nvidia bundle
    Araz Retail API — query stores, inventory, sales, customers, promotions, and orders. CLI: node ~/.openclaw/skills/retail-api/scripts/retail-api.js <command> --telegram-id TID. Authenticate with --telegram-id (the user's Telegram ID from message metadata from.id). Do NOT use --email or --password. Read commands: me, products [--category X] [--brand X] [--season X], inventory [--store N] [--product N] [--low-stock], customers [--id N] [--customer-email X], promotions [--all], sales [--store N], query "SQL". Write commands: transfer --product ID_OR_NAME --to-store N --quantity N [--size S] [--color C] [--from-store N], reorder --product ID_OR_NAME --quantity N [--store N] [--size S] [--color C]. QUERY SYNTAX: query "SELECT ... LIMIT 50" --telegram-id TID. The SQL is a positional argument — do NOT use -q flag (it does not exist). Table names are all lowercase: stores, products, inventory, orders, orderitems, customers, promotions, inventorytransfers, reorderrequests. Products PK is id, name column is product_name
    2.2k repo stars
  39. ▌
    Source Etl Query 2 · nvidia bundle
    Query the host-side source-etls REST mirror for GitHub discussions, historical GitHub mirror data, and NVIDIA forums research.
    2.2k repo stars
  40. ▌
    Github Readonly Live 2 · nvidia bundle
    Read an allowed live GitHub repository through authenticated, policy-scoped GitHub REST GET requests.
    2.2k repo stars
  41. ▌
    Slack Channel Summarizer 2 · nvidia bundle
    Read and summarize Slack channel history from inside the NemoClaw sandbox.
    2.2k repo stars
  42. ▌
    Cross Source Gap Analysis 2 · nvidia
    Compare findings across Slack, GitHub, NVIDIA forums, and Outlook to identify alignment gaps, missing coverage, and follow-ups.
    2.2k repo stars
  43. ▌
    Qiskit To Cudaq · nvidia bundle
    Use when porting Qiskit Python circuits to CUDA-Q kernels while preserving algorithms and validation fidelity.
    2.2k repo stars
  44. ▌
    Nim Operator Install · nvidia bundle
    Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional KServe compatibility verification. Use when a customer wants to install or upgrade the NIM Operator itself, with or without Dynamo and KServe, but does not want to deploy a NIM inference model yet.
    2.2k repo stars
  45. ▌
    Nim Operator Uninstall · nvidia bundle
    Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and post-uninstall validation. Use when a customer wants to remove or clean up the NIM Operator itself, not the GPU Operator or unrelated cluster dependencies.
    2.2k repo stars
  46. ▌
    Cuequivariance · nvidia bundle
    Define custom groups (Irrep subclasses), build segmented tensor products with CG coefficients, create equivariant polynomials and IrDictPolynomials, and use built-in descriptors (linear, tensor products, spherical harmonics). Use when working with cuequivariance group theory, irreps, or segmented polynomials.
    2.2k repo stars
  47. ▌
    Cuequivariance Jax · nvidia bundle
    Execute equivariant polynomials in JAX using segmented_polynomial (naive/uniform_1d), the ir_dict workflow with IrDictPolynomial and dict[Irrep, Array], and Flax NNX layers (IrrepsLinear, SphericalHarmonics, IrrepsIndexedLinear). Use when writing JAX code with cuequivariance.
    2.2k repo stars
  48. ▌
    Cuequivariance Torch · nvidia bundle
    Execute equivariant tensor products in PyTorch using SegmentedPolynomial (naive/uniform_1d/fused_tp/indexed_linear), high-level operations (ChannelWiseTensorProduct, FullyConnectedTensorProduct, Linear, SymmetricContraction, SphericalHarmonics, Rotation), and layers (BatchNorm, FullyConnectedTensorProductConv). Use when writing PyTorch code with cuequivariance.
    2.2k repo stars
  49. ▌
    Nvfleetint · nvidia bundle
    Query NVIDIA Fleet Intelligence with the nvfleetint CLI. Use for ad hoc questions about fleets, nodes, GPUs, node groups, compute zones, alerts, agent health, firmware, verification, inventory, errors, or authentication. For a fleet-wide HTML snapshot use fleet-health-report; for a single-node RCA/RCCA use node-rca-rcca.
    2.2k repo stars
  50. ▌
    Node Rca Rcca · nvidia bundle
    Investigate one NVIDIA Fleet Intelligence node and generate an evidence-backed HTML RCA/RCCA from live current and historical alerts plus authoritative corrective-action research. Use for node incident analysis, root-cause analysis, corrective actions, or post-incident reports.
    2.2k repo stars
  51. ▌
    Fleet Health Report · nvidia bundle
    Generate a standalone fleet-wide HTML health snapshot from live nvfleetint data, including node health, capacity, active-alert impact, recent errors, and machines needing immediate attention. Use for fleet dashboards, executive summaries, or scoped fleet reports. Do not use for a single-node root-cause investigation.
    2.2k repo stars
  52. ▌
    Rebot Orin Setup · nvidia bundle
    Use when initializing a reBot B601-RS (RobStride 6-DOF, seven motors) arm from an NVIDIA Jetson AGX Orin using the Orin's built-in CAN (mttcan) via an external SN65HVD230 transceiver — no USB-CAN/PCAN adapter. Covers the uv/motorbridge env, bringing up can0 @ 1 Mbit, verifying the bus, reboot persistence, motor-ID assignment, zero calibration, and the CAN bring-up troubleshooting tree (loopback-passes-but-no-motors, pinmux, motor power).
    2.2k repo stars
  53. ▌
    Cmake Structure · nvidia
    IsaacTeleop CMake target layout and #include conventions. Use whenever you create or edit a CMakeLists.txt, add or move a module/library/executable/test target, set up target_include_directories or target_link_libraries, decide where a header file goes (private vs public inc/), write or reorder #include directives in C/C++ sources, or restructure directories under src/ or deps/. Also use when reviewing diffs that touch CMakeLists.txt, header placement, include paths, or the overall project/directory structure.
    2.2k repo stars
  54. ▌
    Warp Closing Issue · nvidia bundle
    Use when the user provides Warp commit SHA(s) and GitHub issue number(s) to assess, draft issue comments, post progress updates, or recommend whether issue threads should stay open or close.
    2.2k repo stars
  55. ▌
    Warp Release Audit · nvidia bundle
    Use when generating a Warp pre-release or release-candidate audit report from Towncrier fragments and release history.
    2.2k repo stars
  56. ▌
    Warp Release Notes · nvidia bundle
    Use when drafting GitHub release notes for a Warp feature or bugfix release from Towncrier fragments or a tagged final changelog.
    2.2k repo stars
  57. ▌
    Warp Changelog Audit · nvidia bundle
    Use when auditing and recovering Warp changelog fragments, finalizing a release changelog, or synchronizing a tagged release back to main.
    2.2k repo stars
  58. ▌
    Nvrx Attr · nvidia bundle
    Orchestration layer over nvidia_resiliency_ext attribution modules. Provides log-analysis, fr-analysis, and a Megatron-LM-oriented fault-injection feedback loop for benchmarking attribution quality on SLURM workloads.
    2.2k repo stars
  59. ▌
    Fr Analysis · nvidia bundle
    Analyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available. Detects the wavefront process group where collectives diverge and returns the root-cause suspect ranks.
    2.2k repo stars
  60. ▌
    Log Analysis · nvidia
    Analyze a SLURM job log file for failure root-cause attribution and restart decisions using the packaged NVRx restart-agent. Use when you have a SLURM training job log and need to determine why the job failed and whether it should be restarted. Performs deterministic evidence extraction plus optional LLM enrichment.
    2.2k repo stars
  61. ▌
    Fault Injection Loop · nvidia
    Closed-loop fault injection and attribution accuracy benchmark. Draws from a prioritized pool of (fault_type, rank, iter, nodes) experiments and submits them 2 at a time via sbatch — waiting for each pair to finish before submitting the next — to bound filesystem load. GPU-related faults are front-loaded in the pool. After all jobs complete, runs /log-analysis and /fr-analysis on every experiment, scores attribution vs. ground truth, aggregates gaps, and iterates on attribution modules to close them.
    2.2k repo stars
  62. ▌
    RAG Skill · nvidia bundle
    Retrieve information from simulator manual and example DATA files. Use when answering keyword format questions, syntax queries, or when looking up official documentation and working examples. Essential for understanding keyword definitions, parameter tables, and concrete usage patterns.
    2.2k repo stars
  63. ▌
    Plot Skill · nvidia bundle
    Plot and compare simulation summary metrics. Use when visualizing time-series results, comparing multiple cases, or analyzing production performance. Supports single and multi-metric plots, case comparisons, and automatic metric keyword resolution.
    2.2k repo stars
  64. ▌
    Results Skill · nvidia bundle
    Read and analyze simulation binary output files. Use when extracting summary data, grid properties, or running flow diagnostics (time-of-flight, tracer, allocation, F-Phi, Lorenz) from completed simulations.
    2.2k repo stars
  65. ▌
    Input File Skill · nvidia bundle
    Parse, modify, validate, and patch simulator input files. Use when working with reservoir simulation input files, testing scenarios, or validating simulation configurations. This implementation supports reference format (.DATA); other simulators use different extensions (e.g., .afi, .DAT). Supports natural language modifications, keyword patching, and syntax validation.
    2.2k repo stars
  66. ▌
    Simulation Skill · nvidia bundle
    Run, monitor, and control simulations. Use when executing simulations, checking progress, or stopping running simulations. Supports foreground and background execution, progress monitoring, and process management.
    2.2k repo stars
  67. ▌
    Trt Perf Analysis · nvidia bundle
    Validate and analyze TensorRT performance data from paired layer-info JSON and profile/latency JSON files. Use when asked to inspect TensorRT, TRT, torch-tensorrt, or ONNX-TensorRT perf reports, verify that layer/profile JSON files are valid and from the same model, infer basic model information, find likely fusion or latency optimization opportunities, and produce a concise Markdown performance report or structured JSON data.
    2.2k repo stars
  68. ▌
    Trt Onnx Quickstart · nvidia
    Build and verify a TensorRT engine from a Hugging Face model ID or ONNX file, with numerical parity checked against ONNX Runtime. Use when the user imports a non-LLM model to TensorRT, needs a verified engine from ONNX, hits trtexec "unsupported operator", must verify the engine matches ONNX numerically, debugs a polygraphy parity failure (large max abs diff at FP16), or configures multi-input dynamic shapes. Triggers: convert ONNX to TensorRT, Hugging Face to TensorRT, trtexec onnx, trtexec unsupported operator, optimum-cli export, polygraphy parity check, polygraphy run --trt --onnxrt, parity check failed, max abs diff, verify engine matches ONNX, --minShapes, dynamic shapes trtexec, multi-input shape profile, FP16 engine, INT64 warning. Adjacent skills: `trt-torch-quickstart` (PyTorch frontend), `trt-cpp-runtime-quickstart` (C++ engine load). LLM token generation belongs in TensorRT-LLM, not here.
    2.2k repo stars
  69. ▌
    Trt Torch Quickstart · nvidia
    Compile a PyTorch model to a TensorRT engine via Torch-TensorRT — AOT or JIT — under the new strong-typing default. Use when the user compiles PyTorch to TensorRT without ONNX, hits "enabled_precisions should not be used when use_explicit_typing=True", sees Dynamo graph breaks or PyTorch fallback, debugs ABI errors at import torch_tensorrt, or needs the compatible torch / torch_tensorrt / tensorrt-cu13 version pins for TensorRT 11. Triggers: torch_tensorrt, torch_tensorrt.dynamo.compile, torch.compile backend torch_tensorrt, pytorch to tensorrt, ExportedProgram, Dynamo graph break, use_explicit_typing, enabled_precisions, torch_tensorrt.Input, min_block_size, truncate_double, tensorrt-cu13, version pinning, version compatibility. Adjacent skills: `trt-onnx-quickstart`, `trt-cpp-runtime-quickstart`. LLM token generation belongs in TensorRT-LLM.
    2.2k repo stars
  70. ▌
    Trt Cpp Runtime Quickstart · nvidia
    Load and run a TensorRT engine (.plan / .engine) from C++ using the TensorRT 11 / 10.x **modern Runtime API**, avoiding the deprecated TRT 8.x binding-index APIs that older guidance still promotes. Use whenever the user asks about loading or running a TensorRT .plan/.engine from C++, even on "minimal example" requests — without this skill the default reply uses deprecated enqueueV2-style code. Also use when the user hits "Engine plan file is generated on an incompatible device", deserializeCudaEngine returns nullptr, gets an enqueueV2 / IStreamReader deprecation warning, or wants to stream a .plan via IStreamReaderV2. Triggers: TensorRT C++ inference, load TensorRT plan C++, run .plan from C++, IRuntime example, deserializeCudaEngine, enqueueV3, enqueueV2 deprecated, setTensorAddress, getBindingIndex, IStreamReaderV2, libnvinfer C++. NOT for building engines (`trt-onnx-quickstart`), Python deploy, plugins, multi-GPU.
    2.2k repo stars
  71. ▌
    Trt Strong Typing Migration · nvidia bundle
    Migrate a TensorRT build from weak typing (deprecated 10.12, removed 11.0) to strong typing — across Python INetworkDefinition builders, the trtexec CLI, and C++ builder code. Use when a TRT 11 upgrade breaks a weakly-typed build. Triggers: weakly typed to strongly typed, kSTRONGLY_TYPED, weak typing deprecated, kFP16/kINT8 removed, setPrecision rejected, setComputePrecision deprecated, do I still need --stronglyTyped, how to add the kSTRONGLY_TYPED flag, ModelOpt autocast, INT8 on TRT 11. NOT for ONNX import (`trt-onnx-quickstart`), Torch-TRT (`trt-torch-quickstart`), or C++ deploy (`trt-cpp-runtime-quickstart`).
    2.2k repo stars
  72. ▌
    Deepstream Eval And Finetune · nvidia bundle
    Evaluate and improve an object detector in NVIDIA DeepStream. Use for deployed mAP, FPS and latency measurement, TAO Skill Bank fine-tuning or AutoML, redeployment, and before/after reporting on HuggingFace, NGC, ONNX, or local models and KPI datasets. Object detection only; do not use for classification or non-vision models.
    2.2k repo stars
  73. ▌
    Simple · nvidia bundle
    Summarize short user notes into clear action items.
    2.2k repo stars
  74. ▌
    Task List · nvidia bundle
    Required for 4+ step requests; add tasks at start and update status after each step.
    2.2k repo stars
  75. ▌
    API Caller · nvidia bundle
    Call any REST API dynamically. Make GET, POST, PUT, DELETE requests to any endpoint with custom headers and JSON body.
    2.2k repo stars
  76. ▌
    Calculator · nvidia bundle
    Evaluate mathematical expressions and unit conversions. Handles arithmetic, percentages, exponents, and common unit conversions (temperature, distance, weight). No external dependencies.
    2.2k repo stars
  77. ▌
    Text Analyzer · nvidia bundle
    Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics. Works on any plain text input provided inline or from a file path.
    2.2k repo stars
  78. ▌
    Create Custom Grader · nvidia bundle
    Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
    2.2k repo stars
  79. ▌
    Compileiq Debug · nvidia bundle
    Use when something is wrong: Search() hangs, all evaluations return INVALID_SCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is too high, or a winning ACF candidate needs NCU profiling to explain. Symptom-indexed table on top. Triggers on "compileiq hang", "socket timeout", "INVALID_SCORE", "not converging", "every score is the same", "TypeError fromhex", "ncu profile", "register spill", "ptxas error", "not in expected format", "high cv".
    2.2k repo stars
  80. ▌
    Compileiq Bootstrap · nvidia bundle
    Use when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq-* skill. Verifies CUDA 13.3+, ptxas, GPU access, that `from compileiq.ciq import Search` and friends resolve, and that `PtxasSearchSpace().retrieve()` returns a real path. Documents the env vars that control timeouts, caching, and search-space mirroring. Triggers on "set up compileiq", "compileiq doesn't work", "socket timeout", "where do search spaces come from", "air-gapped compileiq".
    2.2k repo stars
  81. ▌
    Compileiq Run Search · nvidia bundle
    Use when composing the Search(...) call and calling .start(). Covers the four worker classes (MultiProcessWorker / IsoMultiProcessWorker / RayWorker / AsyncWorker) and when to pick each, SearchConfiguration sizing rules, dump_results checkpointing, tracker_config choice (Disabled / Loguru / MLflow), num_workers/task_timeout semantics, and GPU clock locking for stable measurements. Triggers on "Search()", "tuner.start()", "pool_size", "num_workers", "task_timeout", "IsoMultiProcessWorker", "RayWorker", "dump_results", "MLflow", "GPU clocks".
    2.2k repo stars
  82. ▌
    Compileiq Booster Pack · nvidia bundle
    Use BEFORE running a full CompileIQ search. Walks through downloading a Booster Pack from NVIDIA/CompileIQ GitHub Releases, applying ACF candidates one at a time to the user's compiler (raw PTXAS, NVCC, Triton, Helion, FlashInfer), and keeping only candidates that compile, pass correctness, and beat the no-ACF baseline. Includes the mandatory Debug-pack O0/O3 ACF-injection canary that proves the ACF is reaching PTXAS. Triggers on "booster pack", "ACF", "apply-controls", "speed up without searching", "helion fp8", "flashinfer batch decode", "debug pack".
    2.2k repo stars
  83. ▌
    Compileiq Search Space · nvidia
    Use when picking the search_space= argument for Search(). Covers the three provider classes (PtxasSearchSpace, NvccSearchSpace, LocalSearchSpaceBin), how to pin a version/variant/tag, the attention-focused 'att' variant for attention kernels (FlashAttention, GQA, MHA, MLA, FlashInfer Batch Decode), air-gapped mirroring via CIQ_SEARCH_SPACES_DIR, and custom user-defined search spaces built from compileiq.search_spaces.base primitives. Triggers on "search space", "PtxasSearchSpace", "NvccSearchSpace", "air-gapped compileiq", "offline compileiq", "CIQ_SEARCH_SPACES_DIR", "attention variant", "ptxas att".
    2.2k repo stars
  84. ▌
    Compileiq Validate Result · nvidia bundle
    Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF. Loads the dump_results CSV, extracts top-K candidates (single-objective) or the Pareto front (multi-objective), re-measures each against the no-ACF baseline with 100+ trials on fresh caches, runs Welch's t-test plus Cohen's d, rejects three classic false-positive patterns (lucky-min / higher-variance / multiple-comparisons-of-N), and saves the validated winner as best.acf. Triggers on "validate result", "extract best config", "Welch's t-test", "is my speedup real", "save best ACF", "pareto front", "claim speedup", "ship config".
    2.2k repo stars
  85. ▌
    Compileiq Author Objective · nvidia bundle
    Use when writing the objective_function= passed to Search(). Covers the two legal signatures (compiler-only str vs mixed list), the baseline-knockout branch, per-eval cache busting, framework-specific --apply-controls injection (raw PTXAS, NVCC, Triton, Helion, cuTeDSL/FA4, FlashInfer), correctness-before-timing, INVALID_SCORE handling, and the Debug-pack O0/O3 ACF-injection canary that must pass before launching a search. Triggers on "objective function", "apply-controls", "INVALID_SCORE", "save_compiler_config", "baseline knockout", "BASELINE_CONFIG", "every config returns the same score", "TypeError fromhex".
    2.2k repo stars
  86. ▌
    Tripy Testing · nvidia
    Write tests for nvtripy following project conventions. Use when: adding tests for ops, modules, trace operations, or compilation, using pytest parametrize, testing error cases with helper.raises, testing dtype combinations, understanding test directory structure.
    2.2k repo stars
  87. ▌
    Tripy Debugging · nvidia
    Debug and diagnose errors in nvtripy code. Use when: interpreting TripyException stack traces, enabling MLIR/TensorRT debug output, understanding error reporting with stack_info, using raise_error, configuring debug environment variables, tracing compilation failures.
    2.2k repo stars
  88. ▌
    Tripy New Module · nvidia
    Add a new neural network module to nvtripy. Use when: creating an nn layer, implementing a Module subclass, adding a new layer like Linear/LayerNorm/Conv, defining parameters with DefaultParameter or OptionalParameter, using constant_fields decorator.
    2.2k repo stars
  89. ▌
    Tripy Compilation · nvidia
    Work with the nvtripy compilation pipeline. Use when: using tp.compile, creating InputInfo or DimensionInputInfo, understanding the Trace → MLIR → TensorRT flow, configuring optimization levels, working with Executable objects, debugging compilation, using dynamic shapes or NamedDimension.
    2.2k repo stars
  90. ▌
    Tripy Constraints · nvidia
    Author input/output constraints for nvtripy operations using the declarative constraint DSL. Use when: defining input_requirements or output_guarantees, writing @wrappers.interface decorators, auto-casting dtypes, using GetInput/GetReturn/OneOf/If/Equal, debugging constraint validation errors.
    2.2k repo stars
  91. ▌
    Tripy Documentation · nvidia
    Write API documentation for nvtripy following project conventions. Use when: writing docstrings for ops or modules, adding code examples, using @export.public_api document_under paths, creating Sphinx RST cross-references, understanding the docs build pipeline.
    2.2k repo stars
  92. ▌
    Tripy New Operation · nvidia
    Add a new operation to nvtripy. Use when: implementing a new op, adding a frontend op, creating a trace op, registering an op in the API. Covers the full Frontend → Trace → MLIR pipeline including export decorators, constraint definitions, and init registration.
    2.2k repo stars
  93. ▌
    Cosmos3 Setup · nvidia
    Guide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. Use when the user asks "how do I install cosmos3", "how do I set up the environment", "how do I download checkpoints", "how do I use Docker", or any question about getting the package running for the first time.
    2.2k repo stars
  94. ▌
    Cosmos3 Inference · nvidia
    Guide users through running Cosmos3 inference — offline batch generation, online serving with Ray and Gradio, parallelism options, input formats, sampling parameters, and prompt upsampling. Use when the user asks "how do I run inference", "how do I generate a video", "how do I serve the model", "what parameters should I use", or any question about running the model to produce outputs.
    2.2k repo stars
  95. ▌
    Cosmos3 Codebase Nav · nvidia
    Navigate the Cosmos3 package codebase to find where parameters, configs, defaults, scripts, and documentation live. Use when the user asks "where is X in cosmos3", "how do I find the config for Y", "where are the defaults", "where do I change a parameter", or any question about locating files, modules, or settings. Also use when the user opens or edits files and needs orientation.
    2.2k repo stars
  96. ▌
    Cosmos3 Post Training · nvidia
    Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired launch shell recommended, raw `torchrun` as an alternative), running T2V/I2V/V2V inference with the trained DCP checkpoint, and optionally exporting it to Hugging Face safetensors. Use when the user asks how to post-train Cosmos3, fine-tune on a custom video dataset, export a trained checkpoint, or invoke one of the recipe launch shells (`launch_sft_vision_nano.sh`, `launch_sft_llava_ov.sh`, `launch_sft_videophy2_nano.sh`, plus the `_super` LoRA variant) — or any question about `cu130-train` / `cu128-train`, `convert_model_to_dcp` / `export_model` / `train`, or SFT output paths. For dataset captioning / JSONL assembly, see `docs/dataset_jsonl.md`.
    2.2k repo stars
  97. ▌
    Cosmos3 Env Troubleshoot · nvidia
    Diagnose and fix Cosmos3 environment, installation, and runtime errors. Use when the user encounters an ImportError, ModuleNotFoundError, CUDA error, Docker error, checkpoint download failure, or any traceback during setup or inference.
    2.2k repo stars
  98. ▌
    Integrate A Model · nvidia
    End-to-end workflow for porting an external video diffusion model into a flashdreams integration — scope the architecture, scaffold a workspace-member plugin, reuse an existing recipe, write the checkpoint key-remap, layer model-specific conditioners, wire the runner, and verify with checkpoint weight-equality + upstream parity + a GPU rollout. Use when integrating a new model (e.g. a HuggingFace/research release) into flashdreams or a downstream repo, porting upstream weights, or reproducing an existing integration. Pairs with the `flashdreams-integrations` skill (architecture map) — this skill is the ordered procedure; that one is the contract reference.
    2.2k repo stars
  99. ▌
    Maintaining Oss State · nvidia
    Maintain FlashDreams's OSS-release state — the LICENSE / NOTICE / THIRD-PARTY-NOTICES / REUSE.toml / LICENSES/ / CONTRIBUTING.md collateral that satisfies OSRB Bug 6107043, the per-file SPDX headers, the third-party dependency manifest in THIRD-PARTY-NOTICES, and the pyproject.toml + uv.lock dependency pins. Use when adding or upgrading a runtime dependency, vendoring third-party source into the repo, adding a new first-party source file (any .py / .pyx / .pyi / .c / .cc / .cpp / .h / .hpp / .cu / .cuh / .sh / .proto / Dockerfile), reviewing whether a change requires reopening an OSRB bug or filing a self-cert, or triaging a reuse-lint CI failure.
    2.2k repo stars
  100. ▌
    Python Docstring Style · nvidia
    Write Python docstrings and inline comments matching the flashdreams house style — SPDX header, one-line module docstring, Google-style function docstrings (Args/Returns/Raises), PEP 257 attribute docstrings on dataclass/class fields *and on module-level constants*, double-backticks for code references, imperative first sentences, and signpost-style inline block comments (kept, not stripped, on a tightening pass). Use when authoring or editing any .py file under flashdreams/, when adding a new module/class/function/field/constant, when polishing comments, or when the user asks about docstring or comment style.
    2.2k repo stars