Opik Evaluate

Build an LLM evaluation and run it against your app, returning an experiment with scores. Covers datasets, LLM judges, RAG evaluation, synthetic data, error analysis, and validating evaluators against human labels. Use when the user wants to measure or improve AI product quality, or asks about evals, judges, or evaluation metrics.

comet-ml c3661a2 7 files · 47.3 KB Updated

File contents

comet-ml/opik-skills/tree/main/skills/opik-evaluate commit c3661a245e

Frequently asked questions

npx skillmds@latest add comet-ml/opik-evaluate