Llama Cpp 1 0 0

C/C++ LLM inference library with GGUF support, quantization, GPU acceleration (CUDA/Metal/HIP/Vulkan/SYCL), OpenAI-compatible server, and speculative decoding. Use when building local LLM inference applications, deploying models on edge devices, creating OpenAI-compatible API servers, or working with GGUF models.

tangledgroup ebe55c7 8 files · 39.3 KB Updated

File contents

tangledgroup/tangled-skills/tree/main/misc/llama-cpp-1-0-0 commit ebe55c788f

Frequently asked questions

npx skillmds@latest add tangledgroup/llama-cpp-1-0-0