Hqp Sensitivity Aware Hybrid Quantization

Apply the HQP framework to compress and accelerate PyTorch models for edge deployment using sensitivity-aware structural pruning followed by 8-bit post-training quantization. Trigger phrases: 'optimize model for edge', 'prune and quantize model', 'compress model for Jetson', 'reduce inference latency on edge device', 'hybrid quantization and pruning', 'deploy model to edge with size constraints'

ndpvt-web c4b19af 15.8 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/hqp-sensitivity-aware-hybrid-quantization commit c4b19af9da

Frequently asked questions

npx skillmds@latest add ndpvt-web/hqp-sensitivity-aware-hybrid-quantization