Local Inference Tuning

Select and tune a local LLM inference engine for the user's hardware. Use when setting up or auditing local/private model serving, choosing between MLX, llama.cpp, Ollama, and vLLM, estimating model fit, deciding cache/storage policy, tuning batching and KV cache flags, running smoke benchmarks, or exposing an OpenAI-compatible local endpoint.

markoblogo 914d385 3 files · 8.8 KB Updated

File contents

markoblogo/abvx-agent-skills/tree/main/skills/local-inference-tuning commit 914d385df1

Frequently asked questions

npx skillmds@latest add markoblogo/local-inference-tuning