Llama.cpp

Run LLM inference with llama.cpp on CPU, Apple Silicon, AMD/Intel GPUs, or NVIDIA — plus GGUF model conversion and quantization (2–8 bit with K-quants and imatrix). Covers CLI, Python bindings, OpenAI-compatible server, and Ollama/LM Studio integration. Use for edge deployment, M1/M2/M3/M4 Macs, CUDA-less environments, or flexible local quantization.

agentic-in dfaa957 6 files · 40.4 KB Updated

File contents

agentic-in/elephant-agent/tree/main/packages/skills/builtin_packages/mlops/inference/llama-cpp commit dfaa95754d

Frequently asked questions

npx skillmds@latest add agentic-in/llama-cpp