slowlyC
- 4 skills
- 0 followers
- 2 weeks ago last updated
- ▌ Tilelang Skill · slowlyc bundleWrite, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. Use when the task explicitly involves tilelang, tilelang.language, @tilelang.jit, @T.prim_func, T.Kernel, T.copy, T.gemm, TileLang Profiler, Carver, TileLang passes, or TileLang CUDA, ROCm, Metal, and CPU backends. Use triton-skill for Triton or Gluon, cutlass-skill for direct CUTLASS, CuTe, or CuTeDSL work, and cuda-skill for raw CUDA, PTX, NVIDIA architecture, or profiling-tool facts.
- ▌ Cutlass Skill · slowlyc bundleWrite, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Use when the task explicitly involves CUTLASS/CuTe/CuTeDSL, cute::Layout, cute::Tensor, TiledMMA, TiledCopy, CollectiveBuilder, CollectiveMainloop, CollectiveEpilogue, GemmUniversal, KernelSchedule, EpilogueSchedule, CUTLASS pipelines, EVT, pycute, or CUTLASS template errors. Use cuda-skill for raw CUDA/PTX or NVIDIA architecture facts, and triton-skill for Triton or Gluon implementation work.
- ▌ Triton Skill · slowlyc bundleWrite, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl.*, Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, triton_kernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.
- ▌ Cuda Skill · slowlyc bundleQuery current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA, tcgen05, mbarrier, fabric operations, CUDA APIs and Graphs, memory ordering, compute capability, Ampere, Hopper, Blackwell, Rubin, nsys, ncu, and compute-sanitizer.