Optimize Model

Analyze a model server's Python code and its deployment chart to identify performance bottlenecks and optimization opportunities for CPU and GPU (AMD MI300X, ROCm, vLLM) inference services on KServe. Use when you want to improve inference throughput or latency for a model.

wikimedia 0779236 9.2 KB Updated

File contents

wikimedia/machinelearning-liftwing-inference-services/tree/main/.claude/skills/optimize-model commit 0779236321

Frequently asked questions

npx skillmds@latest add wikimedia/optimize-model