Mixed Precision Deployment

Ship FP16, BF16, or FP8 training and inference that holds accuracy while capturing the speedup, using loss scaling and numeric validation. Use when moving a model off FP32 to run faster or fit in less memory and the result must stay correct.

Amey-Thakur c634746 3.5 KB Updated

File contents

Amey-Thakur/AI-SKILLS/tree/main/skills/gpu-ai-infrastructure/mixed-precision-deployment commit c6347463f1

Frequently asked questions

npx skillmds@latest add amey-thakur/mixed-precision-deployment