Turboquant

Implement, use, or explain TurboQuant — Google's data-oblivious vector quantization algorithm for LLM KV cache compression (ICLR 2026). Use this skill when the user asks about KV cache compression, TurboQuant, PolarQuant, QJL (Quantized Johnson-Lindenstrauss), Lloyd-Max quantization for high-dimensional vectors, reducing LLM memory usage, compressing attention keys/values, or implementing any component of the TurboQuant pipeline. Also trigger when the user mentions vector quantization for inference optimization, 3-4 bit KV cache quantization, or inner product preserving compression.

Ryuketsukami Updated

File contents

Ryuketsukami/turboquant-skill/tree/main/ commit 2be3280862

Frequently asked questions

npx skillmds@latest add ryuketsukami/turboquant