Backend
Detection — At the start of every invocation, before taking any action, determine which backend to use:
- If the user passed
--backend pup anywhere in their invocation → use pup mode immediately, regardless of whether MCP tools are present. Skip steps 2–4.
- Check whether MCP tools are present in your active tool list. The canonical signal is whether
mcp__datadog-llmo-mcp__list_llmobs_evals appears in your available tools.
- If MCP tools are present → use MCP mode throughout. Call MCP tools exactly as named in this skill's workflow sections.
- If MCP tools are absent → check whether
pup is executable: run pup --version via Bash. A JSON response containing "version" confirms pup is available.
- If pup responds → use pup mode throughout. Translate every MCP tool call to its pup equivalent using the Tool Reference appendix at the bottom of this file.
- If neither is available → stop and tell the user:
"Neither the Datadog MCP server nor the pup CLI is available. Connect the MCP server (claude mcp add --scope user --transport http datadog-llmo-mcp 'https://mcp.datadoghq.com/api/unstable/mcp-server/mcp?toolsets=llmobs') or install pup."
--backend pup is accepted anywhere in the invocation arguments and is stripped before passing remaining args to the skill logic.
pup invocation rules:
- Invoke via Bash:
pup llm-obs <subcommand> [flags]
- pup always outputs JSON. Parse directly — no content-block unwrapping (unlike MCP results, which may wrap JSON in
[{"type": "text", "text": "<json>"}]).
- If pup returns an auth error, tell the user to run
pup auth login and stop.
- Parallelization: issue multiple Bash tool calls in a single message (one pup command per call).
- Time flags: pup accepts bare duration strings (
1h, 7d, 30m) and RFC3339 timestamps. Do not use now--prefixed strings — strip the prefix when converting from a skill --timeframe argument: now-7d → 7d, now-24h → 24h, now-30d → 30d.
--summary on pup llm-obs spans search strips payload fields to essential metadata only. Use it in bulk/search phases where content is not needed.
Invocation ID: At the very start of each invocation, before any MCP tool call, generate an 8-character hex invocation ID (e.g., 3a9f1c2b). Keep it constant for the entire invocation.
Intent tagging: On every MCP tool call, prefix telemetry.intent with skill:agent-observability-eval-bootstrap[<inv_id>] — followed by a description of why the tool is being called. On the first MCP tool call only, use skill:agent-observability-eval-bootstrap:start[<inv_id>] — instead (note the :start suffix). Example first call: skill:agent-observability-eval-bootstrap:start[3a9f1c2b] — Phase 0: map existing eval coverage for task-cruncher
Eval Bootstrap — Generate Evaluators from Production Traces
Given a sample of production LLM traces, analyze input/output patterns and quality dimensions, then propose a ready-to-use evaluator suite. Four output modes — online evaluators are the default; SDK code, the JSON spec, and the dataset-emit mode are produced on request:
publish (default) — propose online LLM-judge evaluators, then — only after you confirm the suite — write them to Datadog via create_or_update_llmobs_evaluator as disabled drafts (enabled: false). Nothing is created until you confirm at the Phase 2 checkpoint, and nothing scores any spans until you enable it in the UI — the skill never auto-publishes a live evaluator. Once you enable a draft, it runs automatically on matching production spans, traces, or sessions (no dataset, no task function). The skill auto-classifies each proposed evaluator as span-scoped, trace-scoped, or session-scoped based on what the judgment requires (a per-LLM-call tone check vs. an agent goal completion that needs the whole trace vs. user satisfaction across a whole multi-trace conversation) — you accept or override the classification at that checkpoint. Session-scoped evaluators are only proposed when the app's spans carry a session_id (verified by a probe in Phase 1).
sdk_code (on request — --sdk-code, or ask after a publish run) — Python .py file using the Datadog Evals SDK (BaseEvaluator / LLMJudge) for offline experiments.
data_only (on request — --data-only) — self-contained JSON spec, framework-agnostic.
emit_dataset (on request — --emit-dataset <path>) — sample production traces and write a DatasetRecordRaw[] JSON file shaped for LLMObs.create_dataset(records=...). Skips evaluator proposal and generation entirely — this mode produces a dataset, not evaluators. Used by agent-observability-eval-pipeline (Phase 4) to seed an experiment dataset from production behavior.
After a publish run, if the user wants the same suite as offline code or a portable spec, they just ask — the skill regenerates the already-confirmed suite in sdk_code / data_only mode without re-exploring (see "On-request code generation" in Phase 3). The emit_dataset mode is independent of the evaluator workflow and never re-uses a prior proposal — it always re-samples traces.
Usage
/eval-bootstrap <ml_app> [--timeframe <window>] [--sdk-code | --data-only | --emit-dataset <path>] [--trace-limit <N>]
Arguments: $ARGUMENTS
Inputs
| Input |
Required |
Default |
Description |
ml_app |
Yes |
— |
ML application to scope traces |
timeframe |
No |
now-7d |
How far back to look |
rca_report |
No |
— |
Failure taxonomy from eval-trace-rca skill, or a free-text failure hypothesis |
--sdk-code |
No |
off |
Emit a Python SDK .py file for offline experiments instead of publishing online. Mutually exclusive with --data-only and --emit-dataset. |
--data-only |
No |
off |
Emit a self-contained JSON spec file instead of publishing online. Mutually exclusive with --sdk-code and --emit-dataset. |
--emit-dataset <path> |
No |
off |
Dataset-only mode. Sample production traces and write a DatasetRecordRaw[] JSON to <path>. Skips the evaluator workflow entirely. Mutually exclusive with --sdk-code and --data-only. |
--trace-limit |
No |
20 (cap 50) |
Max traces to sample in emit_dataset mode |
If ml_app is missing, ask the user before proceeding. With no mode flag, the skill defaults to publish — it proposes online evaluators and, only after you confirm, creates them as disabled drafts (it never auto-enables them). If more than one of --sdk-code, --data-only, --emit-dataset is supplied, error out and ask which mode the user wants.
Available Tools
| Tool |
Purpose |
search_llmobs_spans |
Find spans by eval presence, tags, span kind, query syntax. Paginate with cursor. |
get_llmobs_span_details |
Metadata, evaluations (scores, labels, reasoning), and content_info map showing available fields + sizes. |
get_llmobs_span_content |
Actual content for a span field. Supports JSONPath via path param for targeted extraction. |
get_llmobs_trace |
Full trace hierarchy as span tree with span counts by kind. |
get_llmobs_agent_loop |
Chronological agent execution timeline (LLM calls, tool invocations, decisions). |
list_llmobs_evals |
List every evaluator configured for the caller's org across all ml_apps, with enabled status and ml_app per result. Call once in Phase 0 to map existing coverage before proposing new evaluators — filter the result by ml_app client-side. |
get_llmobs_evaluator |
Fetch the full persisted evaluator config by name (target ml_app + sampling + filter, provider, prompt template, parsing type, output schema, assessment criteria). Use in Phase 0 to understand what each existing custom eval measures, and (in publish mode) before any update — create_or_update_llmobs_evaluator is full-replace, so you must round-trip the full config to avoid clobbering fields. Not all evaluators have a stored config (notably source=ootb); a not-found error there is expected — skip those. |
create_or_update_llmobs_evaluator |
(publish mode) Write an LLM-judge evaluator config to Datadog. Full-replace semantics: any omitted optional field resets to its default. See "Publishing Conventions" for required fields and structured output → JSON schema mapping. |
delete_llmobs_evaluator |
(publish mode) Only used if the user explicitly asks to remove an evaluator. Never invoke speculatively. |
Key get_llmobs_span_content Patterns
Use the path parameter to extract targeted data without fetching full payloads:
| Field |
Path |
What you get |
messages |
$.messages[0] |
System prompt (first message, usually system role) |
messages |
$.messages[-1] |
Last assistant response |
messages |
(no path) |
Full conversation including tool calls |
input / output |
— |
Span I/O |
documents |
— |
Retrieved documents (RAG apps) |
metadata |
— |
Custom metadata (prompt versions, feature flags, user segments) |
How to Use search_llmobs_spans
Additional filters combine with space (AND): @status:error @ml_app:my-app. Dedicated params (span_kind, root_spans_only, ml_app) work alongside query, but query takes precedence over tags.
To find spans with a specific eval: @evaluations.custom.<eval_name>:* — you can only query for eval presence, not specific results.
To detect whether the app uses sessions: session_id:* matches any span carrying a session_id (session_id is a first-class field — no @ prefix). The Phase 1 session probe uses this to gate session-scope evaluators.
Parallelization Rules
get_llmobs_span_details: Group span_ids by trace_id. One call per trace_id with ALL its span_ids. Issue ALL calls for a page in a single message.
get_llmobs_span_content: Each call is independent — always issue ALL in a single message.
get_llmobs_trace / get_llmobs_agent_loop: Parallelize across different traces in a single message.
- Pipeline parallelism: Start
get_llmobs_span_details for page 1 results immediately — don't wait to collect all pages.
Evaluator SDK Reference
Applies to sdk_code mode only. In data_only mode, use this section as domain context when writing rubric prompts — no SDK classes are emitted.
Imports
# Core classes
from ddtrace.llmobs._experiment import BaseEvaluator, EvaluatorContext, EvaluatorResult
# LLM-as-judge
from ddtrace.llmobs._evaluators.llm_judge import (
LLMJudge,
BooleanStructuredOutput,
ScoreStructuredOutput,
CategoricalStructuredOutput,
)
# Built-in evaluators (use only if needed)
from ddtrace.llmobs._evaluators.format import JSONEvaluator, LengthEvaluator
from ddtrace.llmobs._evaluators.string_matching import StringCheckEvaluator, RegexMatchEvaluator
Only import what the generated file actually uses.
EvaluatorContext (what evaluate() receives)
@dataclass(frozen=True)
class EvaluatorContext:
input_data: dict[str, Any] # Task inputs (from dataset record, NOT from span)
output_data: Any # Task output (from task function return, NOT from span)
expected_output: Optional[JSONType] = None # Ground truth (if available)
metadata: dict[str, Any] = {} # Additional metadata
span_id: Optional[str] = None # LLMObs span ID
trace_id: Optional[str] = None # LLMObs trace ID
Important — span data vs evaluator data: When exploring production traces, you see span I/O (e.g., input.value, output.messages). But evaluators run in offline experiments where input_data and output_data come from the user's dataset records and task function, not from spans. The dataset schema is user-defined and may not match span structure. Write evaluator prompts with generic {{input_data}} / {{output_data}} placeholders and add comments describing what data the evaluator was designed for, so the user can adapt to their dataset shape.
EvaluatorResult (what evaluate() returns)
EvaluatorResult(
value=..., # Required. JSONType (str, int, float, bool, None, list, dict)
reasoning="...", # Optional. Explanation string
assessment="pass" or "fail", # Optional. Pass/fail assessment
metadata={...}, # Optional. Evaluation metadata dict
tags={...}, # Optional. Tags dict
)
LLMJudge — LLM-as-Judge Evaluator
judge = LLMJudge(
user_prompt="...", # Required. Supports {{template_vars}}
system_prompt="...", # Optional. Does NOT support template vars
structured_output=..., # Optional. Boolean/Score/Categorical output, or a dict for custom JSON schema
provider="openai", # "openai" | "anthropic" | "azure_openai" | "vertexai" | "bedrock"
model="gpt-4o", # Model identifier
model_params={"temperature": 0.0}, # Optional. Passed to LLM API
name="eval_name", # Optional. Must match ^[a-zA-Z0-9_-]+$
)
Template variables in user_prompt: {{input_data}}, {{output_data}}, {{expected_output}}, {{metadata.key}} — resolved from EvaluatorContext fields via dot-path into nested dicts.
Structured Output Types
Boolean — true/false with optional pass/fail:
BooleanStructuredOutput(
description="Whether the response is factually accurate",
reasoning=True, # Include reasoning field in LLM response
reasoning_description=None, # Optional custom description for reasoning field
pass_when=True, # True → pass when true, False → pass when false, None → no assessment
)
Score — numeric within a range with optional thresholds:
ScoreStructuredOutput(
description="Helpfulness score",
min_score=1, # Minimum possible score
max_score=10, # Maximum possible score
reasoning=True,
reasoning_description=None,
min_threshold=7, # Scores >= 7 pass (optional)
max_threshold=None, # Scores <= N pass (optional)
)
Categorical — select from predefined categories:
CategoricalStructuredOutput(
categories={
"correct": "The response correctly answers the question",
"partially_correct": "The response is partially correct but missing key information",
"incorrect": "The response is factually wrong or irrelevant",
},
reasoning=True,
reasoning_description=None,
pass_values=["correct"], # Which categories count as passing (optional)
)
Custom JSON schema — arbitrary structured responses for multi-dimensional evals:
# Pass a raw dict as structured_output — used as the JSON schema directly
structured_output={
"type": "object",
"properties": {
"relevance": {"type": "boolean", "description": "Whether the response addresses the question"},
"confidence": {"type": "number", "description": "Confidence score (0.0 to 1.0)"},
"reasoning": {"type": "string", "description": "Explanation for the evaluation"},
},
"required": ["relevance", "confidence", "reasoning"],
"additionalProperties": False,
}
Always write standard JSON schema — the SDK adapts it per provider automatically (e.g., Anthropic doesn't support minimum/maximum on number fields, so the SDK moves range constraints into the description; Vertex AI converts const/anyOf to enum). The full parsed JSON dict becomes the eval value; a "reasoning" key (if present) is automatically extracted. No automatic pass/fail assessment.
LLMJudge Prompt Guidelines
The structured_output parameter enforces the response format via JSON schema. Do not prescribe the format in the prompt (no "Answer YES/NO", "Rate 1-10", etc.). Instead, describe the evaluation criteria and let the structured output handle the format.
- system_prompt: Set the judge's role and the app's domain context. Does NOT support template vars.
- user_prompt: Present the data via
{{input_data}} / {{output_data}}, then describe what good vs. bad looks like for this dimension.
BaseEvaluator — Custom Code-Based Evaluator
For deterministic checks that do not need LLM judgment:
class MyEvaluator(BaseEvaluator):
def __init__(self, name=None, ...custom_params...):
super().__init__(name=name)
self._param = ... # Store config as private attrs
def evaluate(self, context: EvaluatorContext) -> EvaluatorResult:
# Access: context.input_data, context.output_data, context.expected_output, context.metadata
# Must NOT modify self attributes (thread safety)
passed = ... # Your logic here
return EvaluatorResult(
value=passed,
reasoning="...",
assessment="pass" if passed else "fail",
)
Built-in Evaluators
# Validate JSON syntax + optional required keys
JSONEvaluator(required_keys=["name", "age"], output_extractor=None, name=None)
# Validate length (characters, words, or lines)
LengthEvaluator(count_by="words", min_length=10, max_length=500, output_extractor=None, name=None)
# count_by: "characters" | "words" | "lines"
# String matching
StringCheckEvaluator(operation="contains", expected="success", case_sensitive=False, name=None)
# operation: "eq" | "ne" | "contains" | "icontains"
# Regex matching
RegexMatchEvaluator(pattern=r"\d{4}-\d{2}-\d{2}", match_mode="search", name=None)
# match_mode: "search" | "match" | "fullmatch"
Evaluator Type Decision Matrix
| Signal |
Evaluator Type |
| Output must be valid JSON |
JSONEvaluator |
| Output must match a regex pattern |
RegexMatchEvaluator |
| Output has length constraints |
LengthEvaluator |
| Output must contain/not contain specific strings |
StringCheckEvaluator |
| Semantic quality judgment (tone, accuracy, completeness) |
LLMJudge + BooleanStructuredOutput |
| Graded quality on a scale |
LLMJudge + ScoreStructuredOutput |
| Classification into categories |
LLMJudge + CategoricalStructuredOutput |
| Multi-dimensional judgment (evaluate several aspects at once) |
LLMJudge + custom JSON schema dict |
| Complex domain logic combining multiple checks |
BaseEvaluator subclass |
Source Verification
If you have access to dd-trace-py locally, verify the API surface by reading the corresponding modules:
ddtrace.llmobs._evaluators.llm_judge — LLMJudge, BooleanStructuredOutput, ScoreStructuredOutput, CategoricalStructuredOutput
ddtrace.llmobs._experiment — BaseEvaluator, EvaluatorContext, EvaluatorResult
ddtrace.llmobs._evaluators.format — JSONEvaluator, LengthEvaluator
ddtrace.llmobs._evaluators.string_matching — StringCheckEvaluator, RegexMatchEvaluator
Workflow
Phase 0: Resolve Inputs & Entry Mode
Entry mode detection:
| Mode |
Signal |
Behavior |
| Cold Start |
Only ml_app provided (no RCA, no hypothesis) |
Full open discovery — understand what the app does, identify quality dimensions worth measuring, propose evals for coverage |
| From RCA |
Conversation contains an RCA report or user provides a failure hypothesis |
Skip open discovery — use existing failure taxonomy as eval targets |
Parse arguments: Extract ml_app (first non-flag argument), --timeframe (default now-7d), --trace-limit (default 20), --sdk-code, --data-only, and --emit-dataset <path> flags. Set output_mode as follows (at most one of the three mode flags may be set; error if more than one is present):
--emit-dataset <path> set → output_mode = emit_dataset. Skip the rest of the workflow entry-mode logic and jump directly to Phase 3D below.
--sdk-code set → output_mode = sdk_code.
--data-only set → output_mode = data_only.
- otherwise →
output_mode = publish (the default — propose online evaluators, gated on user confirmation, created as disabled drafts).
Resolution steps:
If ml_app not provided → ask the user.
Auto-detect entry mode:
- If the conversation contains an RCA report (look for "Failure Taxonomy" heading, structured failure modes, or severity ratings) →
from_rca. Extract the taxonomy.
- If the user provides a free-text failure hypothesis (e.g., "the system prompt lacks grounding") →
from_rca. Use the hypothesis as the starting eval target.
- Otherwise →
cold_start.
If timeframe not provided → default to now-7d.
Map existing eval coverage — skip if output_mode = data_only (there is no Datadog eval project to check coverage against): Call list_llmobs_evals (org-wide; filter the result client-side to entries where ml_app == <ml_app>). Then, for each eval with source=custom, call get_llmobs_evaluator(eval_name=...) to inspect its prompt template, target, sampling, and filter, and infer which quality dimension it covers. Issue all evaluator calls in a single message (parallelize). Skip source=ootb evals — their names are self-describing and they may not have a fetchable config.
By the end of this step you have a complete coverage map: {eval_name → source, enabled, dimension}. Carry this into Phase 2 for deduplication.
In publish mode, also note any template-variable convention the existing custom evaluators already use (so a new suite reads consistently). Online evaluator templates resolve against the full span JSON, not against EvaluatorContext. See the "Online Template Variables" section under "Publishing Conventions" for the supported syntax ({{span_input}}, {{span_output}}, dot-paths, array selectors, filter accessors).
Notebook context detection: Scan the current conversation for a Datadog notebook URL that was produced by /eval-trace-rca (pattern: https://app.datadoghq.com/notebook/{numeric-id}). If found, store it as rca_notebook_url and extract the numeric ID as rca_notebook_id. This is used after Phase 3 to offer appending the evaluator suite to that notebook instead of creating a new one.
Phase 1: Explore Traces & Identify Eval Targets
Goal: Sample production traces, understand what the app does, and identify quality dimensions worth measuring.
Cold Start Path
Sample the app: search_llmobs_spans(query="@ml_app:\"<ml_app>\" @status:ok", root_spans_only=true, limit=50, from=<timeframe>). Filter by @status:ok — error spans have no output to evaluate.
Session probe (gates session-scope proposals; publish mode): in the same message, also call search_llmobs_spans(query="@ml_app:\"<ml_app>\" session_id:*", limit=20, from=<timeframe>).
- ≥ 1 result → set
sessions_present = true. Note the distinct session_id values and, critically, whether the same session_id appears across multiple trace_ids — that cross-trace span is the real signal that a session carries context worth a session-scope evaluator. (A session_id that only ever maps to one trace adds nothing over trace scope.)
- 0 results →
sessions_present = false. Do not propose any session-scope evaluator; record a one-line "session scope skipped — no session_id on sampled spans" note for the proposal.
Profile the app and identify evaluation target spans: Call get_llmobs_span_details for span_ids grouped by trace_id. Inspect content_info to classify:
| Signal |
App Profile |
content_info has messages |
LLM/chat app |
content_info has documents |
RAG app |
Spans include agent kind |
Agent app |
content_info has metadata |
Has custom metadata |
Multiple span kinds in one trace (agent + tool / retrieval + llm from get_llmobs_trace) |
Multi-step app — at least one trace-scope evaluator likely belongs in the suite (publish mode) |
Same session_id across multiple trace_ids (from the session probe) |
Multi-trace sessions — at least one session-scope evaluator likely belongs in the suite (publish mode, gated on sessions_present) |
For agent/multi-step apps, also call get_llmobs_trace on 2-3 traces to see the full span hierarchy. Compare content_info between the root span and its sub-spans. Then ask two questions for each candidate quality dimension, in this order:
- Does the verdict depend on more than one span? (e.g., faithfulness depends on a
retrieval span's documents AND an llm span's answer; goal completion depends on the chain of tool calls AND the final response.) If yes → trace scope in publish mode. Don't try to compress this into a single span.
- Only if the answer to (1) is no: pick the single span with the richest signal for that dimension (root has the summary; LLM sub-spans have the full system prompt + tool call results + reasoning chain).
Record the span-kind histogram (agent + tool + llm + retrieval) — multiple kinds under one root is a strong signal you'll have at least one trace-scope evaluator in the suite. See Phase 2's "Span vs. Trace vs. Session Scope Classification" for the mandatory walk-through of canonical trace-scope use cases (and, when sessions_present, the canonical session-scope use cases).
Extract content and identify targets: Call get_llmobs_span_content for representative spans. Fetch fields based on app profile:
| App Profile |
Fields to Fetch |
| LLM/chat |
messages (path=$.messages[0] for system prompt), output |
| RAG |
documents, input, output |
| Agent |
get_llmobs_agent_loop for the agent span, then messages for detail |
| Any with metadata |
metadata |
Issue all calls in a single message. As you read, capture two streams of signal:
Generic quality signals — what does "success" look like? What variance exists across outputs? Each observed quality dimension becomes a candidate evaluator, with the traces you've just read as evidence. Also look for safety signals (scope violations, sensitive data in outputs, out-of-character responses) and add a safety evaluator if you find them.
Domain signals — these become the domain-specific evaluator category in Phase 2 (the highest-leverage category). For every 5–10 traces, write down:
- Recurring intents / question categories — what classes of request does this app handle? (
applying for benefit X, comparing flight options, summarizing a policy, creating a widget)
- Entities the app emits in outputs — URLs, agency / company names, code identifiers, monetary amounts, dates, IDs, file paths, phone numbers. Note which ones the user acts on downstream (those are worth a correctness evaluator) versus which are passing references.
- Tool argument shapes (for agent apps) — name each tool the agent calls and the rough schema of its inputs. Tools with non-trivial schemas (≥ 3 fields, structured types) are candidates for argument-correctness evaluators.
- Persona / voice rules — does the app always cite a source, always refuse certain topics (medical, legal, financial advice), always speak in a particular tone? Extract the rules implicitly followed across observed outputs.
- Failure modes specific to the domain — fabricated identifiers, outdated policy references, currency / locale mismatches, off-by-one errors in IDs, wrong units. One observed instance is enough to seed a candidate evaluator.
Don't try to enumerate domain signals exhaustively before reading traces — let the patterns surface as you read. The goal is breadth in the eventual proposal, not completeness in this exploration step.
From RCA Path
Extract the failure taxonomy from the RCA report. Each failure mode with High or Medium severity becomes an eval target. Also run the Phase 1 session probe (query="session_id:*") to set sessions_present — a failure that only manifests across a multi-trace conversation (lost context, repeated mistakes, mounting frustration) is a session-scope target.
Check root cause categories for infrastructure failures. Before proposing evaluators, scan the Root Cause column of the taxonomy for any of: Instrumentation Deficiency, Harness Deficiency, Runtime Error, Upstream Data Issue, or any other root cause that points to infrastructure/environment rather than model behavior. If any are present, pause and ask:
"Some failure modes were diagnosed as infrastructure or instrumentation issues rather than model behavior (e.g., {list the infra root causes}). Evaluators can be designed two ways:
- Behavior-targeted (recommended for ongoing quality): measure whether the model produces correct, specific output — useful once the infrastructure is fixed and you want to track real quality
- Artifact-targeted (useful as regression guard): detect the specific broken output observed (e.g., generic placeholder responses) — catches regressions if the infrastructure breaks again
Which approach do you want, or both?"
- If behavior-targeted: design evaluators for what correct output looks like, not what the broken output looked like. Use the RCA's
expected_output / gold-standard examples as the quality bar.
- If artifact-targeted: design evaluators that detect the specific failure symptom (e.g.,
StringCheckEvaluator for a known bad string, LLMJudge that checks for generic placeholders).
- If both: propose each category separately, clearly labelled.
If all root causes are behavioral (System Prompt Deficiency, Tool Gap, Tool Misuse, Retrieval Failure, etc.) → skip this step and proceed directly.
For each target: if the RCA includes trace IDs, use them directly; otherwise search for matching traces. Fetch 2-3 traces per target with get_llmobs_span_content to understand the concrete pattern.
Phase 2: Propose Evaluator Suite
Goal: Present a concrete evaluator proposal for user confirmation.
In sdk_code / data_only mode — and for eval_scope: span in publish mode — each evaluator judges one data point: input and output for a single record/span, not a full trace or batch. In publish mode, eval_scope: trace judges a whole trace and eval_scope: session a whole multi-trace session — design those against the trace / session payload instead (see "Span vs. Trace vs. Session Scope Classification" below). Design evaluators accordingly for their scope.
Targeting depends on output_mode:
sdk_code / data_only → offline experiments. Template variables use EvaluatorContext fields ({{input_data}}, {{output_data}}). The actual data shape depends on the user's dataset and task function (see EvaluatorContext note in SDK Reference).
publish → online evaluation on production spans. Template variables resolve against the full span JSON via dot-paths ({{meta.input.value}}, {{meta.output.messages[*].content}}, …) or the built-in span-kind-aware aliases ({{span_input}}, {{span_output}}). For eval_scope: trace and eval_scope: session, templates resolve against the trace payload ({{spans[...]}}) or the session payload ({{traces[*].spans[...]}}) instead. See "Online Template Variables" under Publishing Conventions for the full syntax. Each evaluator also needs eval_scope, sampling_percentage, and (optionally) filter — surface these in the proposal table so the user can confirm before publishing. Session scope is only used when the Phase 1 probe set sessions_present.
Order proposals from broadest signal to most granular. Propose broadly, let the user curate — see "How many evaluators to propose" below.
Domain-specific evaluators — What does "good" mean for this specific app? These are the highest-leverage proposals because they capture quality bars generic evaluators miss. Derive them from the domain signals Phase 1 captured:
- Recurring intents / question categories the app handles (e.g., "applying for a federal benefit", "comparing flight options", "explaining a policy"). Propose an
intent_classification or intent_handling_correctness evaluator scoped to the dominant intents.
- Specific entities the app produces (URLs, agency names, code identifiers, monetary amounts, dates, IDs). Propose a per-entity correctness evaluator for the ones with real downstream cost when wrong (e.g.,
cited_url_is_real, agency_name_matches_request, monetary_amount_is_consistent_with_input).
- Tool argument shapes observed across
tool spans. Propose a per-tool argument-correctness evaluator for the tools with non-trivial schemas (e.g., search_flights_args_match_user_request, update_dashboard_widget_targets_correct_widget).
- Persona / voice expectations — does the app always cite sources, always refuse out-of-scope requests, always speak in a specific tone? Propose evaluators for the voice rules you can extract from observed outputs (
cites_a_source, refuses_medical_advice, tone_matches_brand).
- Domain-specific failure modes seen across traces (fabricated identifiers, outdated policy references, unit mismatches, currency / locale mismatches). One evaluator per recurring failure mode.
Name each evaluator after the user-facing concern, not the technical check (agency_url_is_real over regex_url_match). Use the trace IDs you read in Phase 1 as evidence — at least one passing case and one failing case per evaluator if you saw both.
Outcome evaluators — Did this span / trace produce a good result for the request?
- Examples:
task_completion, answer_correctness, response_groundedness
Format evaluators — Does the output meet structural requirements?
- Examples:
valid_json_output, response_length, citation_format
Safety evaluators — Does the output stay within appropriate boundaries?
- Examples:
no_pii_leakage, scope_adherence, no_hallucination
How many evaluators to propose
The default 4-6 cap from the older skill version was too tight — it pushed the skill toward generic evaluators only and left domain signals on the table. Updated guidance:
- Aim for 8–15 evaluators in the proposal, distributed across all four categories (with domain-specific usually the largest bucket, outcome second, format and safety smaller). For very simple single-LLM-call apps, fewer is fine; for agent / RAG apps with rich domain signals, lean toward the upper end.
- Quality > generic: every domain-specific proposal should be backed by at least one observed pattern in the sampled traces. Don't invent generic domain evaluators ("
response_quality") if you don't have evidence for them.
- Let the user curate: the MANDATORY CHECKPOINT below explicitly asks the user to remove what doesn't apply, not just to approve. Treat the proposal as a candidate set the user trims.
Deduplication Against Existing Coverage
In data_only mode: skip this section entirely (coverage map was not built in Phase 0). Proceed directly to the proposal table.
Before building the proposal, apply the coverage map from Phase 0. Coverage is keyed on (dimension, scope) — not on dimension alone: every OOTB evaluator runs at span scope, and an enabled OOTB eval does NOT preclude proposing a trace-scope or session-scope evaluator for the same dimension. The three scopes answer different questions.
Enabled span-scope eval (OOTB or custom) for dimension D:
- Do NOT propose a new span-scope evaluator for D — that dimension is already covered at span scope.
- DO propose a trace-scope or session-scope evaluator for D when the trace or session shape calls for it (multi-step app, or multi-trace session — judgment depends on cross-span or cross-trace context). Note the relationship in the rationale: e.g., "OOTB
Goal Completeness evaluates each LLM span in isolation; this trace-scope goal_completion checks whether the agent's full sequence of steps achieved the user's request, and a session-scope session_goal_completion checks it across the whole conversation — three different questions."
Enabled trace-scope custom eval for dimension D: do NOT propose another trace-scope evaluator for the same dimension; that's a real duplicate. Span-scope on the same dimension is still fair game if the data also fits a single span, and session-scope is fair game if the dimension also needs cross-trace context. Likewise, an enabled session-scope custom eval for D blocks only another session-scope eval for D — span and trace scope remain fair game.
Disabled OOTB eval: Do NOT propose a new custom span-scope evaluator for that dimension. Instead, surface it in a short note within the proposal and suggest enabling it in the Datadog UI rather than creating a duplicate. Example:
hallucination (ootb, disabled) — consider enabling in Datadog UI (Evaluations → Configure) instead of creating a custom span-scope eval. (A trace-scope rag_faithfulness is still in scope and covers a different question.)
Gap identification: Open the proposal with a coverage summary line: "Existing coverage: N evaluator(s) already configured ({names}, all span-scope unless noted). Proposing evaluators for uncovered dimensions and uncovered scopes."
All dimensions covered: A dimension is "fully covered" only when the relevant scopes are present (span, plus trace and/or session where the app shape calls for them). If the coverage map accounts for every identified quality dimension at the appropriate scope(s), surface this explicitly and ask the user what they want: (a) review/improve existing eval prompts, (b) add coverage for additional dimensions, or (c) proceed anyway.
For each proposed evaluator:
- Name: Must match
^[a-zA-Z0-9_-]+$ (alphanumeric, underscore, hyphen only)
- Type:
LLMJudge (Boolean/Score/Categorical/custom JSON schema), built-in (JSONEvaluator, RegexMatchEvaluator, etc.), or BaseEvaluator subclass. In publish mode, only LLM-judge evaluators are supported by the MCP tool — code-based checks must NOT be silently dropped. List them in the same proposal table with Type set to the code-based class, mark them under a "Not publishable in this mode" subsection of the proposal, and tell the user they can get them as offline code on request (--sdk-code, or ask after the publish run) or as a --data-only spec. Treat the code-based proposals as part of the suite for counting and coverage purposes.
- What it measures: 1-2 sentence plain-language description
- Target span: Which span's data the evaluator was designed for (e.g., "root agent span", "LLM sub-span
anthropic.request", "all llm spans"). If the root span's I/O is too lossy for the quality dimension (e.g., tool call results aren't visible), note this and specify which sub-span has the signal. In publish mode this maps to a combination of eval_scope (span/trace/session), root_spans_only, and the EVP filter query (e.g. @meta.span.kind:llm or service:web).
- Pass/fail criteria:
pass_when=True, min_threshold=7, pass_values=["correct"], or "no automatic assessment" for custom JSON schema
- Template variables: Which of
input_data, output_data, expected_output, metadata.* it uses (offline) — or which span paths / aliases it pulls from (publish mode: {{span_input}}, {{span_output}}, {{meta.input.messages[*].content}}, {{meta.metadata.<key>}}, etc.)
- Evidence: At least one trace where it would have caught a failure (or confirmed correct beh
…(truncated)
1---2name: agent-observability-eval-bootstrap3description: Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit Python SDK code or a framework-agnostic JSON spec instead. Use when user says "bootstrap evaluators", "generate evaluators", "create evals from traces", "eval bootstrap", "write evaluators", "build eval suite", "publish evaluators", or wants to generate BaseEvaluator/LLMJudge code or online judge configs from production LLM trace data. Works with ml_app and optional RCA report or failure hypothesis.4---5
6## Backend
7
8**Detection** — At the start of every invocation, before taking any action, determine which backend to use:
9
101. If the user passed `--backend pup` anywhere in their invocation → use **pup mode** immediately, regardless of whether MCP tools are present. Skip steps 2–4.
112. Check whether MCP tools are present in your active tool list. The canonical signal is whether `mcp__datadog-llmo-mcp__list_llmobs_evals` appears in your available tools.
123. If MCP tools are present → use **MCP mode** throughout. Call MCP tools exactly as named in this skill's workflow sections.
134. If MCP tools are absent → check whether `pup` is executable: run `pup --version` via Bash. A JSON response containing `"version"` confirms pup is available.
145. If pup responds → use **pup mode** throughout. Translate every MCP tool call to its pup equivalent using the Tool Reference appendix at the bottom of this file.
156. If neither is available → stop and tell the user:
16 > "Neither the Datadog MCP server nor the pup CLI is available. Connect the MCP server (`claude mcp add --scope user --transport http datadog-llmo-mcp 'https://mcp.datadoghq.com/api/unstable/mcp-server/mcp?toolsets=llmobs'`) or install pup."
17
18`--backend pup` is accepted anywhere in the invocation arguments and is stripped before passing remaining args to the skill logic.
19
20**pup invocation rules:**
21- Invoke via Bash: `pup llm-obs <subcommand> [flags]`
22- pup always outputs JSON. Parse directly — no content-block unwrapping (unlike MCP results, which may wrap JSON in `[{"type": "text", "text": "<json>"}]`).
23- If pup returns an auth error, tell the user to run `pup auth login` and stop.
24- Parallelization: issue multiple Bash tool calls in a single message (one pup command per call).
25- Time flags: pup accepts bare duration strings (`1h`, `7d`, `30m`) and RFC3339 timestamps. Do **not** use `now-`-prefixed strings — strip the prefix when converting from a skill `--timeframe` argument: `now-7d` → `7d`, `now-24h` → `24h`, `now-30d` → `30d`.
26- `--summary` on `pup llm-obs spans search` strips payload fields to essential metadata only. Use it in bulk/search phases where content is not needed.
27
28**Invocation ID:** At the very start of each invocation, before any MCP tool call, generate an 8-character hex invocation ID (e.g., `3a9f1c2b`). Keep it constant for the entire invocation.
29
30**Intent tagging:** On every MCP tool call, prefix `telemetry.intent` with `skill:agent-observability-eval-bootstrap[<inv_id>] — ` followed by a description of why the tool is being called. On the **first MCP tool call only**, use `skill:agent-observability-eval-bootstrap:start[<inv_id>] — ` instead (note the `:start` suffix). Example first call: `skill:agent-observability-eval-bootstrap:start[3a9f1c2b] — Phase 0: map existing eval coverage for task-cruncher`
31
32# Eval Bootstrap — Generate Evaluators from Production Traces
33
34Given a sample of production LLM traces, analyze input/output patterns and quality dimensions, then propose a ready-to-use evaluator suite. Four output modes — **online evaluators are the default**; SDK code, the JSON spec, and the dataset-emit mode are produced on request:
35
36- **`publish`** *(default)* — propose **online** LLM-judge evaluators, then — **only after you confirm the suite** — write them to Datadog via `create_or_update_llmobs_evaluator` as **disabled drafts** (`enabled: false`). Nothing is created until you confirm at the Phase 2 checkpoint, and nothing scores any spans until **you** enable it in the UI — the skill never auto-publishes a live evaluator. Once you enable a draft, it runs automatically on matching production spans, traces, or sessions (no dataset, no task function). The skill **auto-classifies** each proposed evaluator as **span-scoped**, **trace-scoped**, or **session-scoped** based on what the judgment requires (a per-LLM-call tone check vs. an agent goal completion that needs the whole trace vs. user satisfaction across a whole multi-trace conversation) — you accept or override the classification at that checkpoint. Session-scoped evaluators are only proposed when the app's spans carry a `session_id` (verified by a probe in Phase 1).
37- **`sdk_code`** *(on request — `--sdk-code`, or ask after a publish run)* — Python `.py` file using the Datadog Evals SDK (`BaseEvaluator` / `LLMJudge`) for **offline** experiments.
38- **`data_only`** *(on request — `--data-only`)* — self-contained JSON spec, framework-agnostic.
39- **`emit_dataset`** *(on request — `--emit-dataset <path>`)* — sample production traces and write a `DatasetRecordRaw[]` JSON file shaped for `LLMObs.create_dataset(records=...)`. **Skips evaluator proposal and generation entirely** — this mode produces a dataset, not evaluators. Used by `agent-observability-eval-pipeline` (Phase 4) to seed an experiment dataset from production behavior.
40
41After a publish run, if the user wants the same suite as offline code or a portable spec, they just ask — the skill regenerates the **already-confirmed** suite in `sdk_code` / `data_only` mode without re-exploring (see "On-request code generation" in Phase 3). The `emit_dataset` mode is independent of the evaluator workflow and never re-uses a prior proposal — it always re-samples traces.
42
43## Usage
44
45```
46/eval-bootstrap <ml_app> [--timeframe <window>] [--sdk-code | --data-only | --emit-dataset <path>] [--trace-limit <N>]
47```
48
49Arguments: $ARGUMENTS
50
51### Inputs
52
53| Input | Required | Default | Description |
54|-------|----------|---------|-------------|
55| `ml_app` | Yes | — | ML application to scope traces |
56| `timeframe` | No | `now-7d` | How far back to look |
57| `rca_report` | No | — | Failure taxonomy from `eval-trace-rca` skill, or a free-text failure hypothesis |
58| `--sdk-code` | No | off | Emit a Python SDK `.py` file for offline experiments instead of publishing online. Mutually exclusive with `--data-only` and `--emit-dataset`. |
59| `--data-only` | No | off | Emit a self-contained JSON spec file instead of publishing online. Mutually exclusive with `--sdk-code` and `--emit-dataset`. |
60| `--emit-dataset <path>` | No | off | **Dataset-only mode.** Sample production traces and write a `DatasetRecordRaw[]` JSON to `<path>`. Skips the evaluator workflow entirely. Mutually exclusive with `--sdk-code` and `--data-only`. |
61| `--trace-limit` | No | `20` (cap `50`) | Max traces to sample in `emit_dataset` mode |
62
63If `ml_app` is missing, ask the user before proceeding. With no mode flag, the skill defaults to **`publish`** — it proposes online evaluators and, only after you confirm, creates them as disabled drafts (it never auto-enables them). If more than one of `--sdk-code`, `--data-only`, `--emit-dataset` is supplied, error out and ask which mode the user wants.
64
65## Available Tools
66
67| Tool | Purpose |
68|------|---------|
69| `search_llmobs_spans` | Find spans by eval presence, tags, span kind, query syntax. Paginate with cursor. |
70| `get_llmobs_span_details` | Metadata, evaluations (scores, labels, reasoning), and `content_info` map showing available fields + sizes. |
71| `get_llmobs_span_content` | Actual content for a span field. Supports JSONPath via `path` param for targeted extraction. |
72| `get_llmobs_trace` | Full trace hierarchy as span tree with span counts by kind. |
73| `get_llmobs_agent_loop` | Chronological agent execution timeline (LLM calls, tool invocations, decisions). |
74| `list_llmobs_evals` | List every evaluator configured for the caller's org across all ml_apps, with `enabled` status and `ml_app` per result. Call once in Phase 0 to map existing coverage before proposing new evaluators — filter the result by `ml_app` client-side. |
75| `get_llmobs_evaluator` | Fetch the **full** persisted evaluator config by name (target ml_app + sampling + filter, provider, prompt template, parsing type, output schema, assessment criteria). Use in Phase 0 to understand what each existing custom eval measures, and (in publish mode) **before any update** — `create_or_update_llmobs_evaluator` is full-replace, so you must round-trip the full config to avoid clobbering fields. Not all evaluators have a stored config (notably `source=ootb`); a not-found error there is expected — skip those. |
76| `create_or_update_llmobs_evaluator` | *(publish mode)* Write an LLM-judge evaluator config to Datadog. Full-replace semantics: any omitted optional field resets to its default. See "Publishing Conventions" for required fields and structured output → JSON schema mapping. |
77| `delete_llmobs_evaluator` | *(publish mode)* Only used if the user explicitly asks to remove an evaluator. Never invoke speculatively. |
78
79### Key `get_llmobs_span_content` Patterns
80
81Use the `path` parameter to extract targeted data without fetching full payloads:
82
83| Field | Path | What you get |
84|-------|------|-------------|
85| `messages` | `$.messages[0]` | System prompt (first message, usually `system` role) |
86| `messages` | `$.messages[-1]` | Last assistant response |
87| `messages` | *(no path)* | Full conversation including tool calls |
88| `input` / `output` | — | Span I/O |
89| `documents` | — | Retrieved documents (RAG apps) |
90| `metadata` | — | Custom metadata (prompt versions, feature flags, user segments) |
91
92### How to Use `search_llmobs_spans`
93
94Additional filters combine with space (AND): `@status:error @ml_app:my-app`. Dedicated params (`span_kind`, `root_spans_only`, `ml_app`) work alongside `query`, but `query` takes precedence over `tags`.
95
96To find spans with a specific eval: `@evaluations.custom.<eval_name>:*` — you can only query for eval *presence*, not specific results.
97
98To detect whether the app uses **sessions**: `session_id:*` matches any span carrying a `session_id` (`session_id` is a first-class field — no `@` prefix). The Phase 1 session probe uses this to gate session-scope evaluators.
99
100### Parallelization Rules
101
1021. **`get_llmobs_span_details`**: Group span_ids by trace_id. One call per trace_id with ALL its span_ids. Issue ALL calls for a page in a **single message**.
1032. **`get_llmobs_span_content`**: Each call is independent — always issue ALL in a single message.
1043. **`get_llmobs_trace` / `get_llmobs_agent_loop`**: Parallelize across different traces in a single message.
1054. **Pipeline parallelism**: Start `get_llmobs_span_details` for page 1 results immediately — don't wait to collect all pages.
106
107---
108
109## Evaluator SDK Reference
110
111> **Applies to `sdk_code` mode only.** In `data_only` mode, use this section as domain context when writing rubric prompts — no SDK classes are emitted.
112
113### Imports
114
115```python
116# Core classes
117from ddtrace.llmobs._experiment import BaseEvaluator, EvaluatorContext, EvaluatorResult
118
119# LLM-as-judge
120from ddtrace.llmobs._evaluators.llm_judge import (
121 LLMJudge,
122 BooleanStructuredOutput,
123 ScoreStructuredOutput,
124 CategoricalStructuredOutput,
125)
126
127# Built-in evaluators (use only if needed)
128from ddtrace.llmobs._evaluators.format import JSONEvaluator, LengthEvaluator
129from ddtrace.llmobs._evaluators.string_matching import StringCheckEvaluator, RegexMatchEvaluator
130```
131
132Only import what the generated file actually uses.
133
134### EvaluatorContext (what `evaluate()` receives)
135
136```python
137@dataclass(frozen=True)
138class EvaluatorContext:
139 input_data: dict[str, Any] # Task inputs (from dataset record, NOT from span)
140 output_data: Any # Task output (from task function return, NOT from span)
141 expected_output: Optional[JSONType] = None # Ground truth (if available)
142 metadata: dict[str, Any] = {} # Additional metadata
143 span_id: Optional[str] = None # LLMObs span ID
144 trace_id: Optional[str] = None # LLMObs trace ID
145```
146
147**Important — span data vs evaluator data**: When exploring production traces, you see span I/O (e.g., `input.value`, `output.messages`). But evaluators run in offline experiments where `input_data` and `output_data` come from the user's **dataset records and task function**, not from spans. The dataset schema is user-defined and may not match span structure. Write evaluator prompts with generic `{{input_data}}` / `{{output_data}}` placeholders and add comments describing what data the evaluator was designed for, so the user can adapt to their dataset shape.
148
149### EvaluatorResult (what `evaluate()` returns)
150
151```python
152EvaluatorResult(
153 value=..., # Required. JSONType (str, int, float, bool, None, list, dict)
154 reasoning="...", # Optional. Explanation string
155 assessment="pass" or "fail", # Optional. Pass/fail assessment
156 metadata={...}, # Optional. Evaluation metadata dict
157 tags={...}, # Optional. Tags dict
158)
159```
160
161### LLMJudge — LLM-as-Judge Evaluator
162
163```python
164judge = LLMJudge(
165 user_prompt="...", # Required. Supports {{template_vars}}
166 system_prompt="...", # Optional. Does NOT support template vars
167 structured_output=..., # Optional. Boolean/Score/Categorical output, or a dict for custom JSON schema
168 provider="openai", # "openai" | "anthropic" | "azure_openai" | "vertexai" | "bedrock"
169 model="gpt-4o", # Model identifier
170 model_params={"temperature": 0.0}, # Optional. Passed to LLM API
171 name="eval_name", # Optional. Must match ^[a-zA-Z0-9_-]+$
172)
173```
174
175**Template variables** in `user_prompt`: `{{input_data}}`, `{{output_data}}`, `{{expected_output}}`, `{{metadata.key}}` — resolved from `EvaluatorContext` fields via dot-path into nested dicts.
176
177### Structured Output Types
178
179**Boolean** — true/false with optional pass/fail:
180
181```python
182BooleanStructuredOutput(
183 description="Whether the response is factually accurate",
184 reasoning=True, # Include reasoning field in LLM response
185 reasoning_description=None, # Optional custom description for reasoning field
186 pass_when=True, # True → pass when true, False → pass when false, None → no assessment
187)
188```
189
190**Score** — numeric within a range with optional thresholds:
191
192```python
193ScoreStructuredOutput(
194 description="Helpfulness score",
195 min_score=1, # Minimum possible score
196 max_score=10, # Maximum possible score
197 reasoning=True,
198 reasoning_description=None,
199 min_threshold=7, # Scores >= 7 pass (optional)
200 max_threshold=None, # Scores <= N pass (optional)
201)
202```
203
204**Categorical** — select from predefined categories:
205
206```python
207CategoricalStructuredOutput(
208 categories={
209 "correct": "The response correctly answers the question",
210 "partially_correct": "The response is partially correct but missing key information",
211 "incorrect": "The response is factually wrong or irrelevant",
212 },
213 reasoning=True,
214 reasoning_description=None,
215 pass_values=["correct"], # Which categories count as passing (optional)
216)
217```
218
219**Custom JSON schema** — arbitrary structured responses for multi-dimensional evals:
220
221```python
222# Pass a raw dict as structured_output — used as the JSON schema directly
223structured_output={
224 "type": "object",
225 "properties": {
226 "relevance": {"type": "boolean", "description": "Whether the response addresses the question"},
227 "confidence": {"type": "number", "description": "Confidence score (0.0 to 1.0)"},
228 "reasoning": {"type": "string", "description": "Explanation for the evaluation"},
229 },
230 "required": ["relevance", "confidence", "reasoning"],
231 "additionalProperties": False,
232}
233```
234
235Always write standard JSON schema — the SDK adapts it per provider automatically (e.g., Anthropic doesn't support `minimum`/`maximum` on number fields, so the SDK moves range constraints into the `description`; Vertex AI converts `const`/`anyOf` to `enum`). The full parsed JSON dict becomes the eval `value`; a `"reasoning"` key (if present) is automatically extracted. No automatic pass/fail assessment.
236
237### LLMJudge Prompt Guidelines
238
239The `structured_output` parameter enforces the response format via JSON schema. **Do not** prescribe the format in the prompt (no "Answer YES/NO", "Rate 1-10", etc.). Instead, describe the **evaluation criteria** and let the structured output handle the format.
240
241- **system_prompt**: Set the judge's role and the app's domain context. Does NOT support template vars.
242- **user_prompt**: Present the data via `{{input_data}}` / `{{output_data}}`, then describe what good vs. bad looks like for this dimension.
243
244### BaseEvaluator — Custom Code-Based Evaluator
245
246For deterministic checks that do not need LLM judgment:
247
248```python
249class MyEvaluator(BaseEvaluator):
250 def __init__(self, name=None, ...custom_params...):
251 super().__init__(name=name)
252 self._param = ... # Store config as private attrs
253
254 def evaluate(self, context: EvaluatorContext) -> EvaluatorResult:
255 # Access: context.input_data, context.output_data, context.expected_output, context.metadata
256 # Must NOT modify self attributes (thread safety)
257 passed = ... # Your logic here
258 return EvaluatorResult(
259 value=passed,
260 reasoning="...",
261 assessment="pass" if passed else "fail",
262 )
263```
264
265### Built-in Evaluators
266
267```python
268# Validate JSON syntax + optional required keys
269JSONEvaluator(required_keys=["name", "age"], output_extractor=None, name=None)
270
271# Validate length (characters, words, or lines)
272LengthEvaluator(count_by="words", min_length=10, max_length=500, output_extractor=None, name=None)
273# count_by: "characters" | "words" | "lines"
274
275# String matching
276StringCheckEvaluator(operation="contains", expected="success", case_sensitive=False, name=None)
277# operation: "eq" | "ne" | "contains" | "icontains"
278
279# Regex matching
280RegexMatchEvaluator(pattern=r"\d{4}-\d{2}-\d{2}", match_mode="search", name=None)
281# match_mode: "search" | "match" | "fullmatch"
282```
283
284### Evaluator Type Decision Matrix
285
286| Signal | Evaluator Type |
287|--------|---------------|
288| Output must be valid JSON | `JSONEvaluator` |
289| Output must match a regex pattern | `RegexMatchEvaluator` |
290| Output has length constraints | `LengthEvaluator` |
291| Output must contain/not contain specific strings | `StringCheckEvaluator` |
292| Semantic quality judgment (tone, accuracy, completeness) | `LLMJudge` + `BooleanStructuredOutput` |
293| Graded quality on a scale | `LLMJudge` + `ScoreStructuredOutput` |
294| Classification into categories | `LLMJudge` + `CategoricalStructuredOutput` |
295| Multi-dimensional judgment (evaluate several aspects at once) | `LLMJudge` + custom JSON schema `dict` |
296| Complex domain logic combining multiple checks | `BaseEvaluator` subclass |
297
298### Source Verification
299
300If you have access to dd-trace-py locally, verify the API surface by reading the corresponding modules:
301
302- `ddtrace.llmobs._evaluators.llm_judge` — `LLMJudge`, `BooleanStructuredOutput`, `ScoreStructuredOutput`, `CategoricalStructuredOutput`
303- `ddtrace.llmobs._experiment` — `BaseEvaluator`, `EvaluatorContext`, `EvaluatorResult`
304- `ddtrace.llmobs._evaluators.format` — `JSONEvaluator`, `LengthEvaluator`
305- `ddtrace.llmobs._evaluators.string_matching` — `StringCheckEvaluator`, `RegexMatchEvaluator`
306
307---
308
309## Workflow
310
311### Phase 0: Resolve Inputs & Entry Mode
312
313**Entry mode detection:**
314
315| Mode | Signal | Behavior |
316|------|--------|----------|
317| **Cold Start** | Only `ml_app` provided (no RCA, no hypothesis) | Full open discovery — understand what the app does, identify quality dimensions worth measuring, propose evals for coverage |
318| **From RCA** | Conversation contains an RCA report or user provides a failure hypothesis | Skip open discovery — use existing failure taxonomy as eval targets |
319
320**Parse arguments**: Extract `ml_app` (first non-flag argument), `--timeframe` (default `now-7d`), `--trace-limit` (default `20`), `--sdk-code`, `--data-only`, and `--emit-dataset <path>` flags. Set `output_mode` as follows (at most one of the three mode flags may be set; error if more than one is present):
321
322- `--emit-dataset <path>` set → `output_mode = emit_dataset`. Skip the rest of the workflow entry-mode logic and jump directly to **Phase 3D** below.
323- `--sdk-code` set → `output_mode = sdk_code`.
324- `--data-only` set → `output_mode = data_only`.
325- otherwise → `output_mode = publish` (the default — propose online evaluators, gated on user confirmation, created as disabled drafts).
326
327**Resolution steps:**
328
3291. If `ml_app` not provided → ask the user.
3302. Auto-detect entry mode:
331 - If the conversation contains an RCA report (look for "Failure Taxonomy" heading, structured failure modes, or severity ratings) → `from_rca`. Extract the taxonomy.
332 - If the user provides a free-text failure hypothesis (e.g., "the system prompt lacks grounding") → `from_rca`. Use the hypothesis as the starting eval target.
333 - Otherwise → `cold_start`.
3343. If `timeframe` not provided → default to `now-7d`.
3354. **Map existing eval coverage** — **skip if `output_mode = data_only`** (there is no Datadog eval project to check coverage against): Call `list_llmobs_evals` (org-wide; filter the result client-side to entries where `ml_app == <ml_app>`). Then, for each eval with `source=custom`, call `get_llmobs_evaluator(eval_name=...)` to inspect its prompt template, target, sampling, and filter, and infer which quality dimension it covers. Issue all evaluator calls in a **single message** (parallelize). Skip `source=ootb` evals — their names are self-describing and they may not have a fetchable config.
336
337 By the end of this step you have a complete coverage map: `{eval_name → source, enabled, dimension}`. Carry this into Phase 2 for deduplication.
338
339 **In `publish` mode, also note any template-variable convention** the existing custom evaluators already use (so a new suite reads consistently). Online evaluator templates resolve against the **full span JSON**, not against `EvaluatorContext`. See the "Online Template Variables" section under "Publishing Conventions" for the supported syntax (`{{span_input}}`, `{{span_output}}`, dot-paths, array selectors, filter accessors).
340
3415. **Notebook context detection**: Scan the current conversation for a Datadog notebook URL that was produced by `/eval-trace-rca` (pattern: `https://app.datadoghq.com/notebook/{numeric-id}`). If found, store it as `rca_notebook_url` and extract the numeric ID as `rca_notebook_id`. This is used after Phase 3 to offer appending the evaluator suite to that notebook instead of creating a new one.
342
343---
344
345### Phase 1: Explore Traces & Identify Eval Targets
346
347**Goal**: Sample production traces, understand what the app does, and identify quality dimensions worth measuring.
348
349#### Cold Start Path
350
3511. **Sample the app**: `search_llmobs_spans(query="@ml_app:\"<ml_app>\" @status:ok", root_spans_only=true, limit=50, from=<timeframe>)`. Filter by `@status:ok` — error spans have no output to evaluate.
352
353 **Session probe** *(gates session-scope proposals; `publish` mode)*: in the same message, also call `search_llmobs_spans(query="@ml_app:\"<ml_app>\" session_id:*", limit=20, from=<timeframe>)`.
354 - **≥ 1 result** → set `sessions_present = true`. Note the distinct `session_id` values and, critically, whether the same `session_id` appears across **multiple `trace_id`s** — that cross-trace span is the real signal that a session carries context worth a session-scope evaluator. (A `session_id` that only ever maps to one trace adds nothing over trace scope.)
355 - **0 results** → `sessions_present = false`. Do **not** propose any session-scope evaluator; record a one-line "session scope skipped — no `session_id` on sampled spans" note for the proposal.
356
3572. **Profile the app and identify evaluation target spans**: Call `get_llmobs_span_details` for span_ids grouped by trace_id. Inspect `content_info` to classify:
358
359 | Signal | App Profile |
360 |--------|------------|
361 | `content_info` has `messages` | LLM/chat app |
362 | `content_info` has `documents` | RAG app |
363 | Spans include `agent` kind | Agent app |
364 | `content_info` has `metadata` | Has custom metadata |
365 | Multiple span kinds in one trace (`agent` + `tool` / `retrieval` + `llm` from `get_llmobs_trace`) | Multi-step app — at least one trace-scope evaluator likely belongs in the suite (`publish` mode) |
366 | Same `session_id` across **multiple `trace_id`s** (from the session probe) | Multi-trace sessions — at least one session-scope evaluator likely belongs in the suite (`publish` mode, gated on `sessions_present`) |
367
368 For agent/multi-step apps, also call `get_llmobs_trace` on 2-3 traces to see the full span hierarchy. Compare `content_info` between the root span and its sub-spans. Then ask **two** questions for each candidate quality dimension, in this order:
369
370 1. **Does the verdict depend on more than one span?** (e.g., faithfulness depends on a `retrieval` span's documents AND an `llm` span's answer; goal completion depends on the chain of `tool` calls AND the final response.) If yes → **trace scope** in `publish` mode. Don't try to compress this into a single span.
371 2. **Only if the answer to (1) is no**: pick the single span with the richest signal for that dimension (root has the summary; LLM sub-spans have the full system prompt + tool call results + reasoning chain).
372
373 Record the span-kind histogram (agent + tool + llm + retrieval) — multiple kinds under one root is a strong signal you'll have at least one trace-scope evaluator in the suite. See Phase 2's "Span vs. Trace vs. Session Scope Classification" for the mandatory walk-through of canonical trace-scope use cases (and, when `sessions_present`, the canonical session-scope use cases).
374
3753. **Extract content and identify targets**: Call `get_llmobs_span_content` for representative spans. Fetch fields based on app profile:
376
377 | App Profile | Fields to Fetch |
378 |------------|----------------|
379 | LLM/chat | `messages` (`path=$.messages[0]` for system prompt), `output` |
380 | RAG | `documents`, `input`, `output` |
381 | Agent | `get_llmobs_agent_loop` for the agent span, then `messages` for detail |
382 | Any with metadata | `metadata` |
383
384 Issue all calls in a single message. As you read, capture two streams of signal:
385
386 **Generic quality signals** — what does "success" look like? What variance exists across outputs? Each observed quality dimension becomes a candidate evaluator, with the traces you've just read as evidence. Also look for safety signals (scope violations, sensitive data in outputs, out-of-character responses) and add a safety evaluator if you find them.
387
388 **Domain signals** — these become the *domain-specific evaluator* category in Phase 2 (the highest-leverage category). For every 5–10 traces, write down:
389 - **Recurring intents / question categories** — what classes of request does this app handle? (`applying for benefit X`, `comparing flight options`, `summarizing a policy`, `creating a widget`)
390 - **Entities the app emits in outputs** — URLs, agency / company names, code identifiers, monetary amounts, dates, IDs, file paths, phone numbers. Note which ones the user *acts* on downstream (those are worth a correctness evaluator) versus which are passing references.
391 - **Tool argument shapes** (for agent apps) — name each tool the agent calls and the rough schema of its inputs. Tools with non-trivial schemas (≥ 3 fields, structured types) are candidates for argument-correctness evaluators.
392 - **Persona / voice rules** — does the app always cite a source, always refuse certain topics (medical, legal, financial advice), always speak in a particular tone? Extract the rules implicitly followed across observed outputs.
393 - **Failure modes specific to the domain** — fabricated identifiers, outdated policy references, currency / locale mismatches, off-by-one errors in IDs, wrong units. One observed instance is enough to seed a candidate evaluator.
394
395 Don't try to enumerate domain signals exhaustively before reading traces — let the patterns surface as you read. The goal is breadth in the eventual proposal, not completeness in this exploration step.
396
397#### From RCA Path
398
3991. Extract the failure taxonomy from the RCA report. Each failure mode with High or Medium severity becomes an eval target. Also run the Phase 1 **session probe** (`query="session_id:*"`) to set `sessions_present` — a failure that only manifests across a multi-trace conversation (lost context, repeated mistakes, mounting frustration) is a session-scope target.
400
4012. **Check root cause categories for infrastructure failures.** Before proposing evaluators, scan the Root Cause column of the taxonomy for any of: `Instrumentation Deficiency`, `Harness Deficiency`, `Runtime Error`, `Upstream Data Issue`, or any other root cause that points to infrastructure/environment rather than model behavior. If any are present, pause and ask:
402
403 > "Some failure modes were diagnosed as infrastructure or instrumentation issues rather than model behavior (e.g., `{list the infra root causes}`). Evaluators can be designed two ways:
404 > - **Behavior-targeted** (recommended for ongoing quality): measure whether the model produces correct, specific output — useful once the infrastructure is fixed and you want to track real quality
405 > - **Artifact-targeted** (useful as regression guard): detect the specific broken output observed (e.g., generic placeholder responses) — catches regressions if the infrastructure breaks again
406 >
407 > Which approach do you want, or both?"
408
409 - If **behavior-targeted**: design evaluators for what correct output looks like, not what the broken output looked like. Use the RCA's `expected_output` / gold-standard examples as the quality bar.
410 - If **artifact-targeted**: design evaluators that detect the specific failure symptom (e.g., `StringCheckEvaluator` for a known bad string, `LLMJudge` that checks for generic placeholders).
411 - If **both**: propose each category separately, clearly labelled.
412
413 If all root causes are behavioral (System Prompt Deficiency, Tool Gap, Tool Misuse, Retrieval Failure, etc.) → skip this step and proceed directly.
414
4153. For each target: if the RCA includes trace IDs, use them directly; otherwise search for matching traces. Fetch 2-3 traces per target with `get_llmobs_span_content` to understand the concrete pattern.
416
417---
418
419### Phase 2: Propose Evaluator Suite
420
421**Goal**: Present a concrete evaluator proposal for user confirmation.
422
423In `sdk_code` / `data_only` mode — and for `eval_scope: span` in `publish` mode — each evaluator judges **one data point**: input and output for a single record/span, not a full trace or batch. In `publish` mode, `eval_scope: trace` judges a whole trace and `eval_scope: session` a whole multi-trace session — design those against the trace / session payload instead (see "Span vs. Trace vs. Session Scope Classification" below). Design evaluators accordingly for their scope.
424
425**Targeting depends on `output_mode`:**
426
427- `sdk_code` / `data_only` → **offline experiments**. Template variables use `EvaluatorContext` fields (`{{input_data}}`, `{{output_data}}`). The actual data shape depends on the user's dataset and task function (see EvaluatorContext note in SDK Reference).
428- `publish` → **online evaluation on production spans**. Template variables resolve against the **full span JSON** via dot-paths (`{{meta.input.value}}`, `{{meta.output.messages[*].content}}`, …) or the built-in span-kind-aware aliases (`{{span_input}}`, `{{span_output}}`). For `eval_scope: trace` and `eval_scope: session`, templates resolve against the trace payload (`{{spans[...]}}`) or the session payload (`{{traces[*].spans[...]}}`) instead. See "Online Template Variables" under Publishing Conventions for the full syntax. Each evaluator also needs `eval_scope`, `sampling_percentage`, and (optionally) `filter` — surface these in the proposal table so the user can confirm before publishing. Session scope is only used when the Phase 1 probe set `sessions_present`.
429
430Order proposals from broadest signal to most granular. **Propose broadly, let the user curate** — see "How many evaluators to propose" below.
431
4321. **Domain-specific evaluators** — What does "good" mean *for this specific app*? These are the highest-leverage proposals because they capture quality bars generic evaluators miss. Derive them from the **domain signals** Phase 1 captured:
433 - **Recurring intents / question categories** the app handles (e.g., "applying for a federal benefit", "comparing flight options", "explaining a policy"). Propose an `intent_classification` or `intent_handling_correctness` evaluator scoped to the dominant intents.
434 - **Specific entities the app produces** (URLs, agency names, code identifiers, monetary amounts, dates, IDs). Propose a per-entity correctness evaluator for the ones with real downstream cost when wrong (e.g., `cited_url_is_real`, `agency_name_matches_request`, `monetary_amount_is_consistent_with_input`).
435 - **Tool argument shapes** observed across `tool` spans. Propose a per-tool argument-correctness evaluator for the tools with non-trivial schemas (e.g., `search_flights_args_match_user_request`, `update_dashboard_widget_targets_correct_widget`).
436 - **Persona / voice expectations** — does the app always cite sources, always refuse out-of-scope requests, always speak in a specific tone? Propose evaluators for the voice rules you can extract from observed outputs (`cites_a_source`, `refuses_medical_advice`, `tone_matches_brand`).
437 - **Domain-specific failure modes** seen across traces (fabricated identifiers, outdated policy references, unit mismatches, currency / locale mismatches). One evaluator per recurring failure mode.
438
439 Name each evaluator after the *user-facing concern*, not the technical check (`agency_url_is_real` over `regex_url_match`). Use the trace IDs you read in Phase 1 as evidence — at least one passing case and one failing case per evaluator if you saw both.
440
4412. **Outcome evaluators** — Did this span / trace produce a good result for the request?
442 - Examples: `task_completion`, `answer_correctness`, `response_groundedness`
4433. **Format evaluators** — Does the output meet structural requirements?
444 - Examples: `valid_json_output`, `response_length`, `citation_format`
4454. **Safety evaluators** — Does the output stay within appropriate boundaries?
446 - Examples: `no_pii_leakage`, `scope_adherence`, `no_hallucination`
447
448##### How many evaluators to propose
449
450The default `4-6` cap from the older skill version was too tight — it pushed the skill toward generic evaluators only and left domain signals on the table. Updated guidance:
451
452- **Aim for 8–15 evaluators** in the proposal, distributed across all four categories (with domain-specific usually the largest bucket, outcome second, format and safety smaller). For very simple single-LLM-call apps, fewer is fine; for agent / RAG apps with rich domain signals, lean toward the upper end.
453- **Quality > generic**: every domain-specific proposal should be backed by at least one observed pattern in the sampled traces. Don't invent generic domain evaluators ("`response_quality`") if you don't have evidence for them.
454- **Let the user curate**: the MANDATORY CHECKPOINT below explicitly asks the user to **remove** what doesn't apply, not just to approve. Treat the proposal as a candidate set the user trims.
455
456#### Deduplication Against Existing Coverage
457
458**In `data_only` mode**: skip this section entirely (coverage map was not built in Phase 0). Proceed directly to the proposal table.
459
460Before building the proposal, apply the coverage map from Phase 0. **Coverage is keyed on `(dimension, scope)` — not on dimension alone**: every OOTB evaluator runs at span scope, and an enabled OOTB eval does NOT preclude proposing a trace-scope **or session-scope** evaluator for the same dimension. The three scopes answer different questions.
461
4621. **Enabled span-scope eval (OOTB or custom)** for dimension D:
463 - Do NOT propose a new **span-scope** evaluator for D — that dimension is already covered at span scope.
464 - DO propose a **trace-scope** or **session-scope** evaluator for D when the trace or session shape calls for it (multi-step app, or multi-trace session — judgment depends on cross-span or cross-trace context). Note the relationship in the rationale: e.g., "OOTB `Goal Completeness` evaluates each LLM span in isolation; this trace-scope `goal_completion` checks whether the agent's full sequence of steps achieved the user's request, and a session-scope `session_goal_completion` checks it across the whole conversation — three different questions."
465
4662. **Enabled trace-scope custom eval** for dimension D: do NOT propose another trace-scope evaluator for the same dimension; that's a real duplicate. Span-scope on the same dimension is still fair game if the data also fits a single span, and session-scope is fair game if the dimension also needs cross-trace context. Likewise, an enabled **session-scope** custom eval for D blocks only another session-scope eval for D — span and trace scope remain fair game.
467
4683. **Disabled OOTB eval**: Do NOT propose a new custom span-scope evaluator for that dimension. Instead, surface it in a short note within the proposal and suggest enabling it in the Datadog UI rather than creating a duplicate. Example:
469
470 > `hallucination` (ootb, disabled) — consider enabling in Datadog UI (Evaluations → Configure) instead of creating a custom span-scope eval. (A trace-scope `rag_faithfulness` is still in scope and covers a different question.)
471
4724. **Gap identification**: Open the proposal with a coverage summary line: "Existing coverage: N evaluator(s) already configured ({names}, all span-scope unless noted). Proposing evaluators for uncovered dimensions and uncovered scopes."
473
4745. **All dimensions covered**: A dimension is "fully covered" only when the relevant scopes are present (span, plus trace and/or session where the app shape calls for them). If the coverage map accounts for every identified quality dimension at the appropriate scope(s), surface this explicitly and ask the user what they want: (a) review/improve existing eval prompts, (b) add coverage for additional dimensions, or (c) proceed anyway.
475
476For each proposed evaluator:
477
478- **Name**: Must match `^[a-zA-Z0-9_-]+$` (alphanumeric, underscore, hyphen only)
479- **Type**: `LLMJudge` (Boolean/Score/Categorical/custom JSON schema), built-in (`JSONEvaluator`, `RegexMatchEvaluator`, etc.), or `BaseEvaluator` subclass. *In `publish` mode, only LLM-judge evaluators are supported by the MCP tool — code-based checks must NOT be silently dropped. List them in the same proposal table with `Type` set to the code-based class, mark them under a "Not publishable in this mode" subsection of the proposal, and tell the user they can get them as offline code on request (`--sdk-code`, or ask after the publish run) or as a `--data-only` spec. Treat the code-based proposals as part of the suite for counting and coverage purposes.*
480- **What it measures**: 1-2 sentence plain-language description
481- **Target span**: Which span's data the evaluator was designed for (e.g., "root agent span", "LLM sub-span `anthropic.request`", "all `llm` spans"). If the root span's I/O is too lossy for the quality dimension (e.g., tool call results aren't visible), note this and specify which sub-span has the signal. *In `publish` mode this maps to a combination of `eval_scope` (`span`/`trace`/`session`), `root_spans_only`, and the EVP `filter` query (e.g. `@meta.span.kind:llm` or `service:web`).*
482- **Pass/fail criteria**: `pass_when=True`, `min_threshold=7`, `pass_values=["correct"]`, or "no automatic assessment" for custom JSON schema
483- **Template variables**: Which of `input_data`, `output_data`, `expected_output`, `metadata.*` it uses (offline) — or which span paths / aliases it pulls from (publish mode: `{{span_input}}`, `{{span_output}}`, `{{meta.input.messages[*].content}}`, `{{meta.metadata.<key>}}`, etc.)
484- **Evidence**: At least one trace where it would have caught a failure (or confirmed correct beh
485
486…(truncated)