Serve

Compress, deploy, and serve trained ML models in production. Covers model compression (quantization, pruning, distillation, ONNX export), inference APIs, containerization, CI/CD pipelines, monitoring, health endpoints, model versioning, and reproducibility packaging. Use when the user has a trained model and wants to reduce its size, deploy it, serve it, containerize it, build an inference API, set up monitoring, write a model card, create a CI/CD pipeline, or package for reproducibility.

damionrashford c8b9581 5 files · 24.9 KB Updated

File contents

damionrashford/mlx/tree/main/skills/serve commit c8b9581f82

Frequently asked questions

npx skillmds@latest add damionrashford/serve