Fastllm Triton Ops

Guide for adding Triton-backed CUDA operators to FastLLM. Use when modifying FastLLM CUDA op code to add, extend, debug, validate, or benchmark Triton-generated kernels through tools/fastllm_triton_server.py, src/devices/cuda/cudadevice.cpp, src/devices/cuda/fastllm-triton-cuda.cu, include/devices/cuda/fastllm-cuda.cuh, or related CMake wiring.

ztxz16 Updated

File contents

ztxz16/fastllm/tree/main/.codex/skills/fastllm-triton-ops commit e118679fac

Frequently asked questions

npx skillmds@latest add ztxz16/fastllm-triton-ops