Mlx Serving

Use when the user asks about MLX serving, mlx_lm.server, oMLX, Apple Silicon LLM serving, or running a local LLM on a Mac, and when troubleshooting a model that fails to load, OOM during load or inference, a server that hangs or crashes at batch>1, tool calls returned as plaintext content, a throughput regression, or the choice between mlx-lm and oMLX. Also covers oMLX feature-flag tuning (turboquant_kv, dflash, MTP, specprefill, thinking_budget, max-concurrent-requests, force_sampling), the OptiQ proxy for models exceeding RAM, Llama-4 ChunkedKVCache batch handling, Llama-3 tool-call JSON format, and bench-driven validation of serving configs. Apple Silicon M-series only, not cloud LLM hosting, non-MLX backends such as llama.cpp, Ollama and vLLM, or model training.

AeyeOps Updated

File contents

AeyeOps/aeo-skill-marketplace/tree/main/aeo-infra/skills/mlx-serving commit 9d7b8b39c9

Frequently asked questions

npx skillmds@latest add aeyeops/mlx-serving