opik — Open-source LLM Observability, Evaluation & Optimization
Opik (built by Comet) streamlines the
entire LLM application lifecycle: deep tracing of LLM calls and agent
activity, LLM-as-a-judge evaluation, experiment management, production
monitoring at scale (40M+ traces/day), plus the Opik Agent Optimizer and
Opik Guardrails. This skill is the routing-first wrapper — it picks the
right deployment mode, wires the SDK, and drives the trace → evaluate →
monitor → optimize loop.
When to use this skill
- The user asks to install or configure Opik (
pip install opik,
opik configure, ./opik.sh)
- The user wants tracing for LLM calls or agents — via
@opik.track or a
framework integration (OpenAI, Anthropic, LangChain, LangGraph, LlamaIndex,
CrewAI, DSPy, Haystack, Ollama, Bedrock, Vercel AI SDK, Pydantic AI, …)
- The user wants LLM-as-a-judge metrics: hallucination detection, moderation,
answer relevance, context precision/recall for RAG
- The user wants Datasets + Experiments evaluation, or PyTest-integrated
LLM evaluation in CI/CD
- The user wants production LLM monitoring dashboards, online evaluation
rules, prompt playground experiments, agent/prompt optimization, or
guardrails
When not to use this skill
- The stack is LangSmith, not Opik → use
langsmith
- The user needs generic service dashboards/alerts (non-LLM telemetry,
uptime, infra metrics) → use
monitoring-observability
- The user wants offline dataset/KPI interpretation rather than LLM
tracing/eval plumbing → use
data-analysis
- The user is doing root-cause log forensics on app/container logs →
use
log-analysis
Prerequisites
| Requirement |
Notes |
| Python 3.8+ (SDK) |
pip install opik or uv pip install opik |
| Docker + Docker Compose |
Only for local self-hosting via ./opik.sh |
| Kubernetes + Helm |
Only for scalable self-hosted deployments |
| Comet.com account |
Only for the zero-setup cloud option |
Instructions
Step 1 — Choose the server mode
| Mode |
When |
How |
| Comet.com cloud (easiest) |
Quick start, no maintenance |
Create a free account, get API key + workspace |
| Docker Compose (local) |
Local dev & testing, full control |
git clone https://github.com/comet-ml/opik.git && cd opik && ./opik.sh → UI at localhost:5173 |
| Kubernetes / Helm |
Production-scale self-hosting |
Upstream Helm chart guide |
Docker service profiles (development scenarios):
./opik.sh # full Opik suite (default)
./opik.sh --infra # infrastructure only (databases, caches)
./opik.sh --backend # infrastructure + backend services
./opik.sh --guardrails # enable guardrails with any profile
./opik.sh --help # troubleshooting
Windows: powershell -ExecutionPolicy ByPass -c ".\opik.ps1".
Step 2 — Install and configure the SDK
pip install opik # or: uv pip install opik
opik configure # prompts for server address (self-hosted) or API key + workspace (cloud)
Or configure in code:
import opik
opik.configure(use_local=True) # local self-hosted instance
TypeScript, and Ruby (via OpenTelemetry) SDKs are also available — see the
client reference docs.
Step 3 — Log traces
Prefer a native integration when the framework is supported (50+ available:
ADK, AG2, Agno, Anthropic, Autogen, Bedrock, CrewAI, DSPy, Dify, Flowise,
Gemini, Haystack, Instructor, LangChain, LangGraph, LiteLLM, LlamaIndex,
Mastra, Ollama, OpenAI, OpenAI Agents, OpenRouter, OpenTelemetry, Pydantic AI,
Ragas, Semantic Kernel, Smolagents, Spring AI, Vercel AI SDK, WatsonX, …).
See references/commands.md for the full table.
Fallback for any code path — the track decorator (nest-aware, composes
with integrations):
import opik
@opik.track
def my_llm_function(user_question: str) -> str:
# Your LLM code here
return "Hello"
Annotate traces/spans with feedback scores via the SDK or the UI.
Step 4 — Evaluate with LLM-as-a-judge metrics
from opik.evaluation.metrics import Hallucination
metric = Hallucination()
score = metric.score(
input="What is the capital of France?",
output="Paris",
context=["France is a country in Europe."],
)
print(score)
Built-in judges include Hallucination, Moderation, Answer Relevance, Context
Precision/Recall; heuristic metrics and custom metrics are also supported.
Step 5 — Datasets, Experiments, and CI gates
- Manage Datasets and run Experiments to compare prompt/model
variants during development
- Wire evaluations into CI/CD with the PyTest integration so regressions
block merges
- Iterate on prompts/models in the Prompt Playground
Step 6 — Production monitoring and optimization
- Opik is built for scale: 40M+ traces/day ingestion
- Track feedback scores, trace counts, and token usage in the Opik Dashboard
- Add Online Evaluation Rules (LLM-as-a-judge on production traffic) to
catch issues live
- Use Opik Agent Optimizer (dedicated SDK) to improve prompts/agents and
Opik Guardrails for safe-AI policies
Step 7 — Plugin-style installation alongside jeo-skills
This skill folder is plugin-installable through the standard jeo-skills
flow so the wrapper, references, and installer script land on disk for any
supported agent runtime:
# Project install (writes into .agents/skills/opik/)
npx skills add https://github.com/akillness/jeo-skills --skill opik
# Global install for every detected agent
npx skills add -g https://github.com/akillness/jeo-skills --skill opik
# Target specific agents
npx skills add -g https://github.com/akillness/jeo-skills --skill opik -a claude-code -a codex -y
The skill also ships scripts/install.sh as a
one-shot installer covering SDK install (uv → pip fallback) and optional
local self-hosting (OPIK_INSTALL_MODE=local).
Output format
When the user asks opik for help, return a compact brief:
# opik Routing Brief
## Scope
- Server mode: cloud | docker-local | kubernetes | undecided
- SDK: python | typescript | ruby-otel
- Lifecycle stage: tracing | evaluation | ci-gate | production-monitoring | optimization | guardrails
## Recommended next move
- install-sdk | opik-configure | start-local-server | wire-integration | add-judge-metric | create-dataset-experiment | enable-online-rules
## Why
- 2-3 bullets grounded in the user's packet
## Route-outs
- `langsmith` when the observability stack is LangSmith
- `monitoring-observability` for non-LLM dashboards/alerts
- `data-analysis` for offline KPI/metric interpretation
Best practices
- Start with cloud or
./opik.sh, not Kubernetes — Helm is for
production scale; local Docker Compose answers "does tracing work" in
minutes.
- Prefer a native integration over hand-rolled
@opik.track when the
framework is in the support table — integrations capture provider
metadata (tokens, model, latency) automatically.
- Check the changelog before upgrading a self-hosted server — e.g.
v1.7.0 shipped breaking changes.
- Judge metrics need context —
Hallucination and RAG metrics score
against the context you pass; empty context produces misleading scores.
- Gate CI on small, stable datasets — PyTest-integrated experiments
should be fast and deterministic; keep large sweeps in scheduled runs.
- Turn production checks into Online Evaluation Rules instead of
re-running offline experiments against live traffic.
References
1---2name: opik3description: Run Comet's Opik — open-source LLM observability, evaluation, and optimization — from one routing-first skill: install the Python/TypeScript SDK, stand up a server (Comet.com cloud, Docker Compose via `./opik.sh`, or Kubernetes/Helm), wire tracing through `@opik.track` or one of 50+ framework integrations (OpenAI, Anthropic, LangChain, LangGraph, LlamaIndex, CrewAI, DSPy, Ollama, Bedrock, Vercel AI SDK, …), score outputs with LLM-as-a-judge metrics (Hallucination, Moderation, Answer Relevance, Context Precision), and run Datasets/Experiments evaluations including PyTest CI gates. Use when the user wants LLM tracing, prompt evaluation, production LLM monitoring, agent optimization, or guardrails with Opik. Triggers on: opik, comet opik, opik configure, opik.sh, llm observability, llm tracing, llm as a judge, hallucination metric, prompt evaluation, opik dashboard, opik guardrails, agent optimizer.4---56# opik — Open-source LLM Observability, Evaluation & Optimization78[Opik](https://github.com/comet-ml/opik) (built by Comet) streamlines the9entire LLM application lifecycle: deep tracing of LLM calls and agent10activity, LLM-as-a-judge evaluation, experiment management, production11monitoring at scale (40M+ traces/day), plus the **Opik Agent Optimizer** and12**Opik Guardrails**. This skill is the routing-first wrapper — it picks the13right deployment mode, wires the SDK, and drives the trace → evaluate →14monitor → optimize loop.1516## When to use this skill1718- The user asks to install or configure Opik (`pip install opik`,19 `opik configure`, `./opik.sh`)20- The user wants tracing for LLM calls or agents — via `@opik.track` or a21 framework integration (OpenAI, Anthropic, LangChain, LangGraph, LlamaIndex,22 CrewAI, DSPy, Haystack, Ollama, Bedrock, Vercel AI SDK, Pydantic AI, …)23- The user wants LLM-as-a-judge metrics: hallucination detection, moderation,24 answer relevance, context precision/recall for RAG25- The user wants Datasets + Experiments evaluation, or PyTest-integrated26 LLM evaluation in CI/CD27- The user wants production LLM monitoring dashboards, online evaluation28 rules, prompt playground experiments, agent/prompt optimization, or29 guardrails3031## When not to use this skill3233- The stack is **LangSmith**, not Opik → use `langsmith`34- The user needs **generic service dashboards/alerts** (non-LLM telemetry,35 uptime, infra metrics) → use `monitoring-observability`36- The user wants **offline dataset/KPI interpretation** rather than LLM37 tracing/eval plumbing → use `data-analysis`38- The user is doing **root-cause log forensics** on app/container logs →39 use `log-analysis`4041## Prerequisites4243| Requirement | Notes |44|-------------|-------|45| Python 3.8+ (SDK) | `pip install opik` or `uv pip install opik` |46| Docker + Docker Compose | Only for local self-hosting via `./opik.sh` |47| Kubernetes + Helm | Only for scalable self-hosted deployments |48| Comet.com account | Only for the zero-setup cloud option |4950## Instructions5152### Step 1 — Choose the server mode5354| Mode | When | How |55|------|------|-----|56| **Comet.com cloud** (easiest) | Quick start, no maintenance | [Create a free account](https://www.comet.com/signup?from=llm), get API key + workspace |57| **Docker Compose** (local) | Local dev & testing, full control | `git clone https://github.com/comet-ml/opik.git && cd opik && ./opik.sh` → UI at `localhost:5173` |58| **Kubernetes / Helm** | Production-scale self-hosting | Upstream [Helm chart guide](https://www.comet.com/docs/opik/self-host/kubernetes/) |5960Docker service profiles (development scenarios):6162```bash63./opik.sh # full Opik suite (default)64./opik.sh --infra # infrastructure only (databases, caches)65./opik.sh --backend # infrastructure + backend services66./opik.sh --guardrails # enable guardrails with any profile67./opik.sh --help # troubleshooting68```6970Windows: `powershell -ExecutionPolicy ByPass -c ".\opik.ps1"`.7172### Step 2 — Install and configure the SDK7374```bash75pip install opik # or: uv pip install opik76opik configure # prompts for server address (self-hosted) or API key + workspace (cloud)77```7879Or configure in code:8081```python82import opik83opik.configure(use_local=True) # local self-hosted instance84```8586TypeScript, and Ruby (via OpenTelemetry) SDKs are also available — see the87[client reference docs](https://www.comet.com/docs/opik/reference/overview).8889### Step 3 — Log traces9091Prefer a native integration when the framework is supported (50+ available:92ADK, AG2, Agno, Anthropic, Autogen, Bedrock, CrewAI, DSPy, Dify, Flowise,93Gemini, Haystack, Instructor, LangChain, LangGraph, LiteLLM, LlamaIndex,94Mastra, Ollama, OpenAI, OpenAI Agents, OpenRouter, OpenTelemetry, Pydantic AI,95Ragas, Semantic Kernel, Smolagents, Spring AI, Vercel AI SDK, WatsonX, …).96See [`references/commands.md`](references/commands.md) for the full table.9798Fallback for any code path — the `track` decorator (nest-aware, composes99with integrations):100101```python102import opik103104@opik.track105def my_llm_function(user_question: str) -> str:106 # Your LLM code here107 return "Hello"108```109110Annotate traces/spans with feedback scores via the SDK or the UI.111112### Step 4 — Evaluate with LLM-as-a-judge metrics113114```python115from opik.evaluation.metrics import Hallucination116117metric = Hallucination()118score = metric.score(119 input="What is the capital of France?",120 output="Paris",121 context=["France is a country in Europe."],122)123print(score)124```125126Built-in judges include Hallucination, Moderation, Answer Relevance, Context127Precision/Recall; heuristic metrics and custom metrics are also supported.128129### Step 5 — Datasets, Experiments, and CI gates130131- Manage **Datasets** and run **Experiments** to compare prompt/model132 variants during development133- Wire evaluations into CI/CD with the **PyTest integration** so regressions134 block merges135- Iterate on prompts/models in the **Prompt Playground**136137### Step 6 — Production monitoring and optimization138139- Opik is built for scale: 40M+ traces/day ingestion140- Track feedback scores, trace counts, and token usage in the Opik Dashboard141- Add **Online Evaluation Rules** (LLM-as-a-judge on production traffic) to142 catch issues live143- Use **Opik Agent Optimizer** (dedicated SDK) to improve prompts/agents and144 **Opik Guardrails** for safe-AI policies145146### Step 7 — Plugin-style installation alongside jeo-skills147148This skill folder is plugin-installable through the standard jeo-skills149flow so the wrapper, references, and installer script land on disk for any150supported agent runtime:151152```bash153# Project install (writes into .agents/skills/opik/)154npx skills add https://github.com/akillness/jeo-skills --skill opik155156# Global install for every detected agent157npx skills add -g https://github.com/akillness/jeo-skills --skill opik158159# Target specific agents160npx skills add -g https://github.com/akillness/jeo-skills --skill opik -a claude-code -a codex -y161```162163The skill also ships [`scripts/install.sh`](scripts/install.sh) as a164one-shot installer covering SDK install (uv → pip fallback) and optional165local self-hosting (`OPIK_INSTALL_MODE=local`).166167## Output format168169When the user asks `opik` for help, return a compact brief:170171```markdown172# opik Routing Brief173174## Scope175- Server mode: cloud | docker-local | kubernetes | undecided176- SDK: python | typescript | ruby-otel177- Lifecycle stage: tracing | evaluation | ci-gate | production-monitoring | optimization | guardrails178179## Recommended next move180- install-sdk | opik-configure | start-local-server | wire-integration | add-judge-metric | create-dataset-experiment | enable-online-rules181182## Why183- 2-3 bullets grounded in the user's packet184185## Route-outs186- `langsmith` when the observability stack is LangSmith187- `monitoring-observability` for non-LLM dashboards/alerts188- `data-analysis` for offline KPI/metric interpretation189```190191## Best practices1921931. **Start with cloud or `./opik.sh`, not Kubernetes** — Helm is for194 production scale; local Docker Compose answers "does tracing work" in195 minutes.1962. **Prefer a native integration over hand-rolled `@opik.track`** when the197 framework is in the support table — integrations capture provider198 metadata (tokens, model, latency) automatically.1993. **Check the changelog before upgrading a self-hosted server** — e.g.200 v1.7.0 shipped breaking changes.2014. **Judge metrics need context** — `Hallucination` and RAG metrics score202 against the `context` you pass; empty context produces misleading scores.2035. **Gate CI on small, stable datasets** — PyTest-integrated experiments204 should be fast and deterministic; keep large sweeps in scheduled runs.2056. **Turn production checks into Online Evaluation Rules** instead of206 re-running offline experiments against live traffic.207208## References209210- Upstream repo: <https://github.com/comet-ml/opik>211- Documentation: <https://www.comet.com/docs/opik/>212- Quickstart: <https://www.comet.com/docs/opik/quickstart/>213- Integrations overview: <https://www.comet.com/docs/opik/integrations/overview/>214- Metrics overview: <https://www.comet.com/docs/opik/evaluation/metrics/overview/>215- Self-host (local): <https://www.comet.com/docs/opik/self-host/local_deployment>216- Self-host (Kubernetes): <https://www.comet.com/docs/opik/self-host/kubernetes/>217- Installer script: [`scripts/install.sh`](scripts/install.sh)218- Command + integration reference: [`references/commands.md`](references/commands.md)219- Adjacent skills: `../langsmith/SKILL.md`, `../monitoring-observability/SKILL.md`,220 `../data-analysis/SKILL.md`, `../log-analysis/SKILL.md`221- License: Apache-2.0 (see upstream `LICENSE`)