title: "Ollama"
description: "Run large language models locally with a simple REST API and CLI. Supports Llama, Mistral, Gemma, Qwen, DeepSeek, and 100+ models. Zero-config setup, model library, multimodal support, and OpenAI-compatible API. Best for local development, offline inference, and privacy-sensitive workloads."
skillName: "ollama"
skillVersion: "1.0.0"
skillAuthor: "Orchestra Research"
skillLicense: "MIT"
skillTags: ["Inference", "Ollama", "Local LLMs", "REST API", "OpenAI Compatible", "Privacy", "Offline", "Llama", "Mistral", "Gemma"]
skillDeps: ["ollama", "httpx", "openai"]
|
|
| Version |
1.0.0 |
| Author |
Orchestra Research |
| License |
MIT |
| Tags |
Inference Ollama Local LLMs REST API OpenAI Compatible Privacy Offline |
| Dependencies |
ollama httpx openai |
Ollama — Run LLMs Locally
The simplest way to run open-source LLMs on your own hardware.
When to use Ollama
Use Ollama when:
- Local development without cloud API costs
- Privacy-sensitive workloads (data never leaves your machine)
- Offline inference required
- Rapid prototyping with 100+ available models
- OpenAI API drop-in replacement needed
Metrics:
- 100,000+ GitHub stars
- 100+ supported models (Llama 3, Mistral, Gemma, Qwen, DeepSeek, Phi)
- Zero-config setup — single binary
Quick start
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.2
ollama run llama3.2
Python SDK
import ollama
response = ollama.chat(
model='llama3.2',
messages=[{'role': 'user', 'content': 'What is GRPO training?'}]
)
print(response['message']['content'])
# Streaming
for chunk in ollama.generate(model='llama3.2', prompt='Explain LoRA', stream=True):
print(chunk['response'], end='', flush=True)
OpenAI-compatible drop-in
from openai import OpenAI
client = OpenAI(base_url='http://localhost:11434/v1', api_key='ollama')
response = client.chat.completions.create(
model='llama3.2',
messages=[{'role': 'user', 'content': 'Explain attention mechanisms'}]
)
Embeddings for RAG
import ollama
response = ollama.embeddings(model='nomic-embed-text', prompt='Flash Attention paper')
embedding = response['embedding'] # 768-dim vector
Custom Modelfile
FROM llama3.2
SYSTEM "You are an expert AI researcher. Provide precise, citation-backed answers."
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
ollama create research-assistant -f Modelfile
ollama run research-assistant
Model library
| Model |
Size |
Best for |
| llama3.2 |
3B, 8B |
General purpose |
| llama3.1:70b |
70B |
High quality |
| mistral |
7B |
Instruction following |
| qwen2.5-coder |
7B, 32B |
Code generation |
| deepseek-r1 |
8B–70B |
Reasoning |
| phi4 |
14B |
Compact + capable |
| llava |
7B, 13B |
Vision + language |
REST API
curl http://localhost:11434/api/generate -d '{"model":"llama3.2","prompt":"Explain FSDP","stream":false}'
References
1---2name: ollama3description: ---4---5---6 title: "Ollama"7 description: "Run large language models locally with a simple REST API and CLI. Supports Llama, Mistral, Gemma, Qwen, DeepSeek, and 100+ models. Zero-config setup, model library, multimodal support, and OpenAI-compatible API. Best for local development, offline inference, and privacy-sensitive workloads."8 skillName: "ollama"9 skillVersion: "1.0.0"10 skillAuthor: "Orchestra Research"11 skillLicense: "MIT"12 skillTags: ["Inference", "Ollama", "Local LLMs", "REST API", "OpenAI Compatible", "Privacy", "Offline", "Llama", "Mistral", "Gemma"]13 skillDeps: ["ollama", "httpx", "openai"]14 ---1516 | | |17 |---|---|18 | **Version** | 1.0.0 |19 | **Author** | Orchestra Research |20 | **License** | MIT |21 | **Tags** | `Inference` `Ollama` `Local LLMs` `REST API` `OpenAI Compatible` `Privacy` `Offline` |22 | **Dependencies** | `ollama` `httpx` `openai` |232425 # Ollama — Run LLMs Locally2627 The simplest way to run open-source LLMs on your own hardware.2829 ## When to use Ollama3031 **Use Ollama when:**32 - Local development without cloud API costs33 - Privacy-sensitive workloads (data never leaves your machine)34 - Offline inference required35 - Rapid prototyping with 100+ available models36 - OpenAI API drop-in replacement needed3738 **Metrics**:39 - **100,000+ GitHub stars**40 - **100+ supported models** (Llama 3, Mistral, Gemma, Qwen, DeepSeek, Phi)41 - **Zero-config setup** — single binary4243 ## Quick start4445 ```bash46 curl -fsSL https://ollama.com/install.sh | sh47 ollama pull llama3.248 ollama run llama3.249 ```5051 ### Python SDK5253 ```python54 import ollama5556 response = ollama.chat(57 model='llama3.2',58 messages=[{'role': 'user', 'content': 'What is GRPO training?'}]59 )60 print(response['message']['content'])6162 # Streaming63 for chunk in ollama.generate(model='llama3.2', prompt='Explain LoRA', stream=True):64 print(chunk['response'], end='', flush=True)65 ```6667 ### OpenAI-compatible drop-in6869 ```python70 from openai import OpenAI71 client = OpenAI(base_url='http://localhost:11434/v1', api_key='ollama')72 response = client.chat.completions.create(73 model='llama3.2',74 messages=[{'role': 'user', 'content': 'Explain attention mechanisms'}]75 )76 ```7778 ### Embeddings for RAG7980 ```python81 import ollama82 response = ollama.embeddings(model='nomic-embed-text', prompt='Flash Attention paper')83 embedding = response['embedding'] # 768-dim vector84 ```8586 ### Custom Modelfile8788 ```dockerfile89 FROM llama3.290 SYSTEM "You are an expert AI researcher. Provide precise, citation-backed answers."91 PARAMETER temperature 0.292 PARAMETER num_ctx 819293 ```9495 ```bash96 ollama create research-assistant -f Modelfile97 ollama run research-assistant98 ```99100 ## Model library101102 | Model | Size | Best for |103 |-------|------|----------|104 | llama3.2 | 3B, 8B | General purpose |105 | llama3.1:70b | 70B | High quality |106 | mistral | 7B | Instruction following |107 | qwen2.5-coder | 7B, 32B | Code generation |108 | deepseek-r1 | 8B–70B | Reasoning |109 | phi4 | 14B | Compact + capable |110 | llava | 7B, 13B | Vision + language |111112 ## REST API113114 ```bash115 curl http://localhost:11434/api/generate -d '{"model":"llama3.2","prompt":"Explain FSDP","stream":false}'116 ```117118 ## References119 - [Ollama GitHub](https://github.com/ollama/ollama)120 - [Model Library](https://ollama.com/library)121 - [API Reference](https://github.com/ollama/ollama/blob/main/docs/api.md)122