Ollama
Run LLMs locally with simple commands and a REST API.
Quick Start
ollama pull llama3.2
ollama run llama3.2 "What is the capital of France?"
When to Use
- Private/local LLM inference
- Offline AI applications
- Testing models without API costs
- Custom fine-tuned models
Step-by-Step
- Install Ollama from ollama.com
- Pull a model:
ollama pull llama3.2 - Run interactively or via API
- Create custom Modelfiles
Dependencies
# Install from https://ollama.com
ollama pull llama3.2:3b
Examples
import requests
response = requests.post("http://localhost:11434/api/generate", json={
"model": "llama3.2",
"prompt": "Why is the sky blue?",
"stream": False
})
print(response.json()["response"])
Resources
Validation
- Ollama service is running
- Model pulls and runs successfully
- API returns responses at localhost:11434