SlowlyC@Agent Gpu Skills

SlowlyC@Agent Gpu Skills from gabrielmoreira/agent-skills-mirror.

by @gabrielmoreira 4 skills

Skills in this plugin

4
  1. Cuda Skill · slowlyc bundle
    Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA, tcgen05, mbarrier, fabric operations, CUDA APIs and Graphs, memory ordering, compute capability, Ampere, Hopper, Blackwell, Rubin, nsys, ncu, and compute-sanitizer.
    2 installs
  2. Triton Skill · slowlyc bundle
    Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl.*, Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, triton_kernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.
    1 install
  3. Cutlass Skill · slowlyc bundle
    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Use when the task explicitly involves CUTLASS/CuTe/CuTeDSL, cute::Layout, cute::Tensor, TiledMMA, TiledCopy, CollectiveBuilder, CollectiveMainloop, CollectiveEpilogue, GemmUniversal, KernelSchedule, EpilogueSchedule, CUTLASS pipelines, EVT, pycute, or CUTLASS template errors. Use cuda-skill for raw CUDA/PTX or NVIDIA architecture facts, and triton-skill for Triton or Gluon implementation work.
    1 install
  4. Tilelang Skill · slowlyc bundle
    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. Use when the task explicitly involves tilelang, tilelang.language, @tilelang.jit, @T.prim_func, T.Kernel, T.copy, T.gemm, TileLang Profiler, Carver, TileLang passes, or TileLang CUDA, ROCm, Metal, and CPU backends. Use triton-skill for Triton or Gluon, cutlass-skill for direct CUTLASS, CuTe, or CuTeDSL work, and cuda-skill for raw CUDA, PTX, NVIDIA architecture, or profiling-tool facts.
    1 install