# Debug With Langwatch

> Root-cause production errors and misbehaving agent runs with LangWatch. Finds errored traces, inspects spans, checks monitor and evaluator scores, then narrows to a root cause. Use when something is failing or misbehaving in production (errors, bad answers, latency spikes).

- Skill: `langwatch/debug-with-langwatch-3` (Agent Skill)
- Install (CLI): `npx skillmds@latest add langwatch/debug-with-langwatch-3`
- Raw SKILL.md: https://api.skillmd.com/api/skills/langwatch/debug-with-langwatch-3/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- License: MIT
- Author: langwatch (https://skillmd.com/u/langwatch)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/langwatch/debug-with-langwatch-3

---


# Debug Production Issues with LangWatch

A structured diagnostic workflow: errored traces → span inspection → monitor/evaluator scores → root cause. Work the steps in order; each narrows the search space for the next.

If traces themselves look broken (empty inputs/outputs, disconnected spans), switch to the `debug-instrumentation` recipe instead. That is an instrumentation problem, not an application problem.

## Prerequisites

Use `langwatch docs <path>` to read documentation as Markdown. Some useful entry points:

```bash
langwatch docs                                    # Docs index
langwatch docs integration/python/guide           # Python integration
langwatch docs integration/typescript/guide       # TypeScript integration
langwatch docs prompt-management/cli              # Prompts CLI
langwatch scenario-docs                           # Scenario docs index
```

Discover commands with `langwatch --help` and `langwatch <subcommand> --help`. List and get commands accept `--format json` for machine-readable output. Every list command takes `--limit <n>` to cap the rows and `--jq <expr>` to read part of the answer. A paginated list answers with an envelope, so count its rows through the row array (`--jq '.traces | length'`), and read how many there are in all at `.pagination.total`. Bare `--jq length` counts the fields of the envelope, not the rows. Read the docs first instead of guessing SDK APIs or CLI flags.

If no shell is available, fetch the same Markdown over plain HTTP. Append `.md` to any docs path (e.g. https://langwatch.ai/docs/integration/python/guide.md). Index: https://langwatch.ai/docs/llms.txt. Scenario index: https://langwatch.ai/scenario/llms.txt

If anything fails or confuses you while following this skill (broken commands, docs that do not match reality, errors you had to work around), ask the user for permission and run `npx --yes langwatch report --user-approved` with a `--title` and `--summary` (or `--session <transcript.jsonl>`) to send it to the LangWatch team, and it directly shapes what gets fixed. No login or API key needed. Nothing is sent without `--user-approved`, and `--dry-run` prints the exact payload without sending anything. The title, summary and transcript are scrubbed locally first, by pattern: secrets and API keys, plus email addresses, phone numbers, card numbers and public IPv4 addresses. Anything no pattern matches is sent as written, including a contact address passed with `--email`. With `--session`, always run `--dry-run` first and let the user read the payload, because a transcript carries content they never reviewed. `npx --yes langwatch report --help` explains the options.

## Step 0: Point the CLI at the Right Project

```bash
langwatch status
```

A fast sanity check that the API key, endpoint, and project are the ones you mean to debug. Fix auth first (see the `setup-lw` recipe): every later step reads from this project.

## Step 1: Find the Errored Traces

```bash
langwatch trace search --errors-only --limit 25 -o json
langwatch trace search --errors-only -q "timeout" --start-date 2026-01-01 -o json
```

- `--errors-only` is how you find failures. An error is recorded on the span, not in the trace's searchable text, so `-q "error"` finds nothing and reads like a clean project.
- `--start-date`/`--end-date` bound the window (ISO strings or epoch ms; default is the last 24h).
- `-q` does a text search over one phrase: the error message, a user id, a thread id. AND, OR and NOT are matched as words, not as operators.
- The result is `{ "traces": [...], "pagination": { "totalHits": N } }`. Pull fields out with `--jq` instead of reading the whole payload:

```bash
langwatch trace search --limit 50 -o json --jq ".traces[].traceId"
langwatch trace search -q "refund" -o json --jq ".traces | length"
```

Look for: traces with error statuses, empty or truncated outputs, outliers in latency or cost, and repeats of the same failure across users/threads (a pattern, not a one-off).

## Step 2: Inspect the Failing Spans

```bash
langwatch trace get <traceId>            # human-readable digest
langwatch trace get <traceId> -o json    # full span hierarchy
```

Read the span tree top-down:

- **Which span failed?** The error is usually in one span (an LLM call, a tool call), not the whole trace. Note its input: a bad input upstream often explains a failure downstream.
- **What did the model see?** Check the prompt/messages on the failing LLM span. Missing context, truncated history, and stale retrieved documents are the usual suspects.
- **Retries and timeouts:** repeated identical spans suggest retry loops; a long-running span before the failure suggests a timeout.

## Step 3: Check Monitors and Evaluator Scores

Production quality signals live in monitors (online evaluation) and their evaluators:

```bash
langwatch monitor list -o json           # which monitors exist, are they enabled/firing?
langwatch monitor get <id> -o json       # one monitor's config and recent state
langwatch evaluator list -o json         # the evaluators the monitors run
```

- A firing monitor names the failure mode (toxicity, hallucination, PII). Corroborate it against the spans from Step 2.
- No monitor for the failure mode you found? That is a gap worth closing once the root cause is fixed (`langwatch monitor create`).

For a quantitative view of the blast radius:

```bash
langwatch analytics query -m trace-count -a sum --group-by metadata.model -o json
```

## Step 4: Root Cause and Verify

1. Form a hypothesis from the failing span's input + the monitor's failure mode: prompt change, model change, bad retrieval, code regression. `git log` on the agent's code and prompts tells you what changed when the failures started.
2. Apply the fix (prompt, code, or configuration).
3. Generate fresh traffic, then re-run Step 1: the errored traces should stop appearing.
4. If the failure was a regression, add a scenario so it stays fixed. The `scenarios` skill covers this.

## Discovery

The full command surface, with per-command usage hints, is one command away:

```bash
langwatch commands -o json     # machine-readable catalog of every command
langwatch help-tree            # compact annotated tree (fits in context)
langwatch <group> --help       # flags for one group
```

Pass `--agent` to any command for compact single-line JSON with colour and spinners off (the CLI sets it by itself under Claude Code, Cursor, Copilot CLI and Amazon Q).

