Llama Cpp

Runs local GGUF inference with llama.cpp (llama-cli, llama-server, optional llama-cpp-python) on CPU, Metal, CUDA, ROCm, or Intel GPU, including Hugging Face -hf Hub downloads. Use when the user wants llama.cpp, GGUF files, llama-cli, or llama-server. Not for Ollama install/serve (ollama-local-setup) or Ollama Cloud GLM (ollama). Never treat mmproj-*.gguf projector files as the main weights.

Kayforkind baa65dd 7 files · 44.7 KB Updated

File contents

Kayforkind/skill-slice commit baa65dd3e1

Frequently asked questions

npx skillmds@latest add kayforkind/llama-cpp