Optimize Local Model Compression

Use when a developer needs to compress a local model for tool-calling workloads — "which quantization should I use", "why does my model emit broken JSON", "how do I get better tool-call fidelity from a 4-bit model", "OptiQ vs QAT vs naive", "the model stopped calling tools after quantization". Covers layer-aware sensitivity-driven compression, QAT group-size matching, and outcome-optimized calibration for structured output.

understudylabs Updated

File contents

understudylabs/understudy-agent-tools/tree/main/skills/optimize-local-model-compression commit 7743048afe

Frequently asked questions

npx skillmds@latest add understudylabs/optimize-local-model-compression