Accuracy Safe Quantization

Shrink a converted LiteRT model with ai-edge-quantizer (fp16 / int8 / int4) without losing accuracy, verifying parity against the float source after every step. Use when choosing a quantization recipe for a new model, when a quantized model fails to load, degrades on a task benchmark, or degenerates over long generations, or when deciding between dynamic-range, weight-only, and blockwise variants.

google-ai-edge 14363e4 10.3 KB Updated

File contents

google-ai-edge/litert-samples/tree/main/skills/accuracy-safe-quantization commit 14363e4ba4

Frequently asked questions

npx skillmds@latest add google-ai-edge/accuracy-safe-quantization