Llama Cpp

LLM inference in C/C++ with Python bindings. GPU acceleration via CUDA/Metal/Vulkan, 2-8 bit quantization (GGUF), KV cache, and grammar-based sampling. Run Llama, Mistral, Gemma, Phi locally.

mkurman ddd9a90 1.5 KB Updated

File contents

mkurman/zorai/tree/main/skills/scientific-skills/llama-cpp commit ddd9a90d39

Frequently asked questions

npx skillmds@latest add mkurman/llama-cpp