Hqq Quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.

tianhao909 Updated 1 repo stars

File contents

tianhao909/AI-Research-SKILLs-cn/tree/main/10-optimization/hqq commit 4a73b9b1a5

Frequently asked questions

npx skillmds@latest add tianhao909/hqq-quantization