DeepEval Tracing
Use this skill to instrument an AI application — an LLM app, agent, RAG
pipeline, or chatbot — with DeepEval's native tracing so its execution is
visible span by span in Confident AI's Observatory. The work is: pick a
supported integration when one exists, fall back to manual @observe
otherwise, give each span a meaningful type, and add tags and metadata.
This skill stops at producing well-formed traces. Attaching evaluation metrics
and running evals is the deepeval skill's job.
Scope: AI Applications Only
Instrument only the AI parts of the system — agent loops and planning, LLM
calls, retrieval / vector search, and tool calls. The span types (llm,
retriever, tool, agent) describe AI components. Do not trace non-AI
software (web servers, CRUD backends, infrastructure). If the target has no
LLM, agent, retrieval, or tool-calling component, this skill does not apply.
When to Use vs the deepeval and deepeval-otel Skills
- This skill (
deepeval-tracing) — instrument an app with the DeepEval SDK
(@observe, framework integrations) so traces reach Confident AI.
deepeval skill — build pytest eval suites: datasets, metrics, traced
evals, deepeval test run, iteration. It runs evals against an app this
skill instrumented.
deepeval-otel skill — instrument with the vendor-neutral OpenTelemetry
SDK instead of the DeepEval SDK (raw OTLP, including non-Python apps).
The three are complementary. If unsure between this skill and deepeval-otel:
use this one when the app is Python and you want the DeepEval SDK; use
deepeval-otel when you want raw OpenTelemetry or the app is not Python.
Prerequisites
- An AI application in Python with
pip install deepeval.
- For traces to reach Confident AI:
deepeval login, or an exported
CONFIDENT_API_KEY (preferred for CI and non-interactive runs).
Workflow
- Confirm the target is an AI application (it has LLM calls, an agent loop,
retrieval, or tool calls). If it has none of these, stop — this skill does
not apply.
- Detect the framework, model provider, agent SDK, and vector database in use.
- Read
references/integrations.md and the exact integration doc for what was
detected. Prefer a native integration over manual instrumentation.
- If no native integration fits, instrument manually with
@observe. Read
references/tracing.md.
- Give each span a meaningful
type (llm, retriever, tool, agent) and
capture inputs/outputs.
- Add trace-level tags and metadata where they help diagnose failure patterns.
Never trace secrets, credentials, or raw sensitive data.
- Confirm
deepeval login or CONFIDENT_API_KEY, then verify traces appear
in the Confident AI Observatory.
Core Principles
- Instrument AI components only —
llm, retriever, tool, agent spans.
Never trace non-AI software.
- Prefer a supported integration over manual
@observe. Manual tracing is the
fallback for unsupported frameworks and app-owned wrapper boundaries.
- Read the exact integration doc before writing tracing code.
- Give spans meaningful types; let names default to function names unless
there is a strong reason to override.
- Never trace secrets, credentials, API keys, or raw sensitive user data.
- Producing traces is the scope. Attaching metrics and running evals belong to
the
deepeval skill; raw OpenTelemetry export belongs to deepeval-otel.
References
| Topic |
File |
Manual instrumentation: @observe, span types, tags, metadata |
references/tracing.md |
| Integration selection rule and framework / model / vector-DB doc index |
references/integrations.md |
Source: confident-ai/deepeval — distributed by TomeVault.
1---2name: confident-ai-deepeval-deepeval3description: DeepEval Tracing4---56# DeepEval Tracing78Use this skill to instrument an **AI application** — an LLM app, agent, RAG9pipeline, or chatbot — with **DeepEval's native tracing** so its execution is10visible span by span in **Confident AI's Observatory**. The work is: pick a11supported integration when one exists, fall back to manual `@observe`12otherwise, give each span a meaningful type, and add tags and metadata.1314This skill stops at producing well-formed traces. Attaching evaluation metrics15and running evals is the `deepeval` skill's job.1617## Scope: AI Applications Only1819Instrument only the AI parts of the system — agent loops and planning, LLM20calls, retrieval / vector search, and tool calls. The span types (`llm`,21`retriever`, `tool`, `agent`) describe AI components. Do not trace non-AI22software (web servers, CRUD backends, infrastructure). If the target has no23LLM, agent, retrieval, or tool-calling component, this skill does not apply.2425## When to Use vs the `deepeval` and `deepeval-otel` Skills2627- **This skill (`deepeval-tracing`)** — instrument an app with the DeepEval SDK28 (`@observe`, framework integrations) so traces reach Confident AI.29- **`deepeval` skill** — build pytest eval suites: datasets, metrics, traced30 evals, `deepeval test run`, iteration. It runs evals *against* an app this31 skill instrumented.32- **`deepeval-otel` skill** — instrument with the vendor-neutral OpenTelemetry33 SDK instead of the DeepEval SDK (raw OTLP, including non-Python apps).3435The three are complementary. If unsure between this skill and `deepeval-otel`:36use this one when the app is Python and you want the DeepEval SDK; use37`deepeval-otel` when you want raw OpenTelemetry or the app is not Python.3839## Prerequisites4041- An AI application in Python with `pip install deepeval`.42- For traces to reach Confident AI: `deepeval login`, or an exported43 `CONFIDENT_API_KEY` (preferred for CI and non-interactive runs).4445## Workflow46471. Confirm the target is an AI application (it has LLM calls, an agent loop,48 retrieval, or tool calls). If it has none of these, stop — this skill does49 not apply.502. Detect the framework, model provider, agent SDK, and vector database in use.513. Read `references/integrations.md` and the exact integration doc for what was52 detected. Prefer a native integration over manual instrumentation.534. If no native integration fits, instrument manually with `@observe`. Read54 `references/tracing.md`.555. Give each span a meaningful `type` (`llm`, `retriever`, `tool`, `agent`) and56 capture inputs/outputs.576. Add trace-level tags and metadata where they help diagnose failure patterns.58 Never trace secrets, credentials, or raw sensitive data.597. Confirm `deepeval login` or `CONFIDENT_API_KEY`, then verify traces appear60 in the Confident AI Observatory.6162## Core Principles63641. Instrument AI components only — `llm`, `retriever`, `tool`, `agent` spans.65 Never trace non-AI software.662. Prefer a supported integration over manual `@observe`. Manual tracing is the67 fallback for unsupported frameworks and app-owned wrapper boundaries.683. Read the exact integration doc before writing tracing code.694. Give spans meaningful types; let names default to function names unless70 there is a strong reason to override.715. Never trace secrets, credentials, API keys, or raw sensitive user data.726. Producing traces is the scope. Attaching metrics and running evals belong to73 the `deepeval` skill; raw OpenTelemetry export belongs to `deepeval-otel`.7475## References7677| Topic | File |78| --- | --- |79| Manual instrumentation: `@observe`, span types, tags, metadata | `references/tracing.md` |80| Integration selection rule and framework / model / vector-DB doc index | `references/integrations.md` |8182---83> Source: [confident-ai/deepeval](https://github.com/confident-ai/deepeval) — distributed by [TomeVault](https://tomevault.io).84<!-- tomevault:4.0:skill_md:2026-06-24 -->