llama.cpp Portable LLM Inference Engine in C/C++

llama.cpp is a high-performance C/C++ implementation for running LLM inference across diverse hardware. It supports GGUF model quantization, GPU acceleration on NVIDIA/AMD/Apple Silicon, and provides both a CLI and an OpenAI-compatible HTTP server for local model serving.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/llama-cpp-portable-llm-inference commit bd13632c7f

Frequently asked questions

npx skillmds@latest add agentskillexchange/llama-cpp-portable-llm-inference-engine-in-c-c