Hqq Quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.

qcmuu Updated 0 repo stars

File contents

qcmuu/AI-Research-Skills/tree/main/10-optimization/hqq commit 4a73b9b1a5

Frequently asked questions

npx skillmds@latest add qcmuu/hqq-quantization