Llama Cpp Windows Deployment · kan-h bundle llama.cpp Windows 平台多模型部署与优化。涵盖:Router Mode 多模型管理、Gemma 4/Qwen/Phi-4 模型部署、QAT–MTP 推理优化、RTX 5060 Ti 16GB 显存调优、CPU 内存受限场景适配、预编译包快速部署、Preset 差异化配置、WSL2 通路配置、Agent/API 接入、Tool Calling 工具调用测试、纯 CPU 工具调用脚本、BAT 脚本编码修复。Use when: 部署 llama.cpp 服务、配置多模型路由、优化推理性能、排查显存/内存溢出、配置 MTP 投机解码、迁移 Qwen/Gemma 模型、测试模型工具调用、搭建纯 CPU 推理、修复 .bat 闪退/乱码。English: deploying llama.cpp on Windows, configuring Router Mode, optimizing GPU/CPU inference, troubleshooting OOM, setting up MTP speculative decoding, migrating between Gemma 4 and Qwen models, testing tool calling, running CPU-only inference, fixing .bat encoding crashes.