Gguf Quantization

Convert and quantize models to GGUF format for efficient CPU/GPU inference with llama.cpp, supporting 2-8 bit quantization and Apple Silicon acceleration.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/10-optimization/gguf commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/gguf-quantization