# llama.cpp Portable LLM Inference Engine in C/C++

> llama.cpp is a high-performance C/C++ implementation for running LLM inference across diverse hardware. It supports GGUF model quantization, GPU acceleration on NVIDIA/AMD/Apple Silicon, and provides both a CLI and an OpenAI-compatible HTTP server for local model serving.

- Skill: `agentskillexchange/llama-cpp-portable-llm-inference-engine-in-c-c` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/llama-cpp-portable-llm-inference-engine-in-c-c`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/llama-cpp-portable-llm-inference-engine-in-c-c/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/llama-cpp-portable-llm-inference-engine-in-c-c

---


# llama.cpp Portable LLM Inference Engine in C/C++

llama.cpp is a high-performance C/C++ implementation for running LLM inference across diverse hardware. It supports GGUF model quantization, GPU acceleration on NVIDIA/AMD/Apple Silicon, and provides both a CLI and an OpenAI-compatible HTTP server for local model serving.

## Installation

Use the upstream install or setup path that matches your environment:
- Run with Docker - see our [Docker documentation](docs/docker.md)

Requirements and caveats from upstream:
- Python: [ddh0/easy-llama](https://github.com/ddh0/easy-llama)
- Python: [abetlen/llama-cpp-python](https://github.com/abetlen/llama-cpp-python)
- Node.js: [withcatai/node-llama-cpp](https://github.com/withcatai/node-llama-cpp)

Basic usage or getting-started notes:
- Install llama.cpp using [brew, nix or winget](docs/install.md)
- Download pre-built binaries from the [releases page](https://github.com/ggml-org/llama.cpp/releases)
- Build from source by cloning this repository - check out [our build guide](docs/build.md)

- Source: https://github.com/ggml-org/llama.cpp
- Extracted from upstream docs: https://raw.githubusercontent.com/ggml-org/llama.cpp/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/llama-cpp-portable-llm-inference/)

