# Debug With Langwatch

> Root-cause production errors and misbehaving agent runs with LangWatch. Finds errored traces, inspects spans, checks monitor and evaluator scores, then narrows to a root cause. Use when something is failing or misbehaving in production (errors, bad answers, latency spikes).

- Skill: `langwatch/debug-with-langwatch` (Agent Skill)
- Install (CLI): `npx skillmds@latest add langwatch/debug-with-langwatch`
- Raw SKILL.md: https://api.skillmd.com/api/skills/langwatch/debug-with-langwatch/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- License: MIT
- Author: langwatch (https://skillmd.com/u/langwatch)
- Updated: 2026-09-10
- Page: https://skillmd.com/skills/langwatch/debug-with-langwatch

---


# Debug Production Issues with LangWatch

A structured diagnostic workflow: errored traces → span inspection → monitor/evaluator scores → root cause. Work the steps in order; each narrows the search space for the next.

If traces themselves look broken (empty inputs/outputs, disconnected spans), switch to the `debug-instrumentation` recipe instead. That is an instrumentation problem, not an application problem.

## Prerequisites

## Step 0: Point the CLI at the Right Project

```bash
langwatch status
```

A fast sanity check that the API key, endpoint, and project are the ones you mean to debug. Fix auth first (see the `setup-lw` recipe): every later step reads from this project.

## Step 1: Find the Errored Traces

```bash
langwatch trace search --errors-only --limit 25 -o json
langwatch trace search --errors-only -q "timeout" --start-date 2026-01-01 -o json
```

- `--errors-only` is how you find failures. An error is recorded on the span, not in the trace's searchable text, so `-q "error"` finds nothing and reads like a clean project.
- `--start-date`/`--end-date` bound the window (ISO strings or epoch ms; default is the last 24h).
- `-q` does a text search over one phrase: the error message, a user id, a thread id. AND, OR and NOT are matched as words, not as operators.
- The result is `{ "traces": [...], "pagination": { "totalHits": N } }`. Pull fields out with `--jq` instead of reading the whole payload:

```bash
langwatch trace search --limit 50 -o json --jq ".traces[].traceId"
langwatch trace search -q "refund" -o json --jq ".traces | length"
```

Look for: traces with error statuses, empty or truncated outputs, outliers in latency or cost, and repeats of the same failure across users/threads (a pattern, not a one-off).

## Step 2: Inspect the Failing Spans

```bash
langwatch trace get <traceId>            # human-readable digest
langwatch trace get <traceId> -o json    # full span hierarchy
```

Read the span tree top-down:

- **Which span failed?** The error is usually in one span (an LLM call, a tool call), not the whole trace. Note its input: a bad input upstream often explains a failure downstream.
- **What did the model see?** Check the prompt/messages on the failing LLM span. Missing context, truncated history, and stale retrieved documents are the usual suspects.
- **Retries and timeouts:** repeated identical spans suggest retry loops; a long-running span before the failure suggests a timeout.

## Step 3: Check Monitors and Evaluator Scores

Production quality signals live in monitors (online evaluation) and their evaluators:

```bash
langwatch monitor list -o json           # which monitors exist, are they enabled/firing?
langwatch monitor get <id> -o json       # one monitor's config and recent state
langwatch evaluator list -o json         # the evaluators the monitors run
```

- A firing monitor names the failure mode (toxicity, hallucination, PII). Corroborate it against the spans from Step 2.
- No monitor for the failure mode you found? That is a gap worth closing once the root cause is fixed (`langwatch monitor create`).

For a quantitative view of the blast radius:

```bash
langwatch analytics query -m trace-count -a sum --group-by metadata.model -o json
```

## Step 4: Root Cause and Verify

1. Form a hypothesis from the failing span's input + the monitor's failure mode: prompt change, model change, bad retrieval, code regression. `git log` on the agent's code and prompts tells you what changed when the failures started.
2. Apply the fix (prompt, code, or configuration).
3. Generate fresh traffic, then re-run Step 1: the errored traces should stop appearing.
4. If the failure was a regression, add a scenario so it stays fixed. The `scenarios` skill covers this.

## Discovery

The full command surface, with per-command usage hints, is one command away:

```bash
langwatch commands -o json     # machine-readable catalog of every command
langwatch help-tree            # compact annotated tree (fits in context)
langwatch <group> --help       # flags for one group
```

Pass `--agent` to any command for compact single-line JSON with colour and spinners off (the CLI sets it by itself under Claude Code, Cursor, Copilot CLI and Amazon Q).

