Debug E2E Test
Start from a test name (not a CI run), find its recent distinct failure modes,
investigate one, falsify a root-cause mechanism, and land on fix-the-test vs.
file-a-bug with the action to match. This is an orchestrator: deterministic work
lives in scripts/, and detailed procedures live in references/ that you read
only when a stage needs them.
Editing this skill: this file holds the invariant, the reference holds the mechanism. Restating a mechanism here makes two copies that drift apart silently.
When to use
You picked up one specific e2e test that is failing or flaking, and want its evidence without hunting for it by hand. Two entries, one triage. They differ only in where the evidence comes from; everything from "read the summary" onward is identical.
| CI entry | Local entry | |
|---|---|---|
| Use when | the test fails/flakes in CI | it just failed on this machine |
| Evidence | test-health history -> one occurrence's report |
this run's test-results/ artifacts |
| Needs | E2E_INSIGHTS_API_KEY, gh |
nothing |
| Extra steps | pattern table, which-pattern question, prior-triage check | none |
Pick local when the engineer just produced the failure or says "locally"; pick CI
when they name a test CI is failing. When they said neither, don't ask -- look:
run collect-local-evidence.js --test '<what they named>' (drop --test if they
named nothing), and take the local entry if it returns ok (that call is the
local entry's first step), the CI entry otherwise. They compose -- a local dig that needs a rate runs the history
query afterwards, and a CI diagnosis reproduces locally in the verification half.
This is the debugging process for this case. Don't layer a general debugging workflow on top of it -- the rules below (one pattern at a time, a falsifiable mechanism, evidence ruled in and out) are that discipline. If you arrived here mid-way through another one, drop it and restart from the evidence; a hypothesis formed before the evidence was read is the thing this skill exists to prevent.
Not this skill: a test still being written or edited (author-e2e-tests); a
whole run's failures or a run ID/URL (e2e-failure-analyzer); a Vitest or
extension-host failure.
Non-negotiable rules
These hold on both entries unless a line names one.
- Zero runs is never a clean result (CI) -- only nonzero runs with no failure patterns is. No local artifacts is not a dead end (local) -- it means the test hasn't run yet, so it ends in an offer to run it.
- Resolve the test identity with
resolve-test-key.js, never by hand; when it returns candidates instead of a resolution, ask which test before querying. The local entry needs this only to build a run command. - Investigate one selected pattern at a time (CI); ask which when there's more than one. Never fetch evidence for a pattern the engineer didn't select -- ask first, even to check a side theory about how patterns relate.
- Never run a test in the foreground. Use Bash
run_in_background: trueand tail the log. A foreground call lasting minutes blocks delivery of the engineer's messages until it returns, so "it's hung, stop it" cannot reach you mid-run and you cannotTaskStopthe run. This matters most here: debugging means watching a run you may well want to abandon. - Fetch one representative occurrence first; a second only for a listed
reason in
references/evidence-escalation.md-- name which. - Agree the fix approach before the first edit, the same way you agree the pattern before fetching evidence.
- Escalate evidence only to answer a concrete question, and show the evidence
block before you act on it. Large output stays on disk and out of the
conversation; how the read is delegated is
references/evidence-escalation.md. - Never increase a timeout or add an arbitrary wait as the fix.
- Never claim a flaky test is fixed on one green run.
- A previous merged fix must be checked against subsequent failures, and that
check reported as four lines, not a triage report (
references/prior-triage.md). - Root-cause claims cite observed evidence and the alternatives ruled out.
- Checkpoint at every phase transition (CI). The local entry starts one only
once it escalates (
references/local-evidence.md).
Requirements
- Claude Code, run from a Positron checkout: the scripts resolve the repo root from their own location and keep triage state in the shared git dir.
- On PATH:
nodeandgit. The CI entry additionally needsgh(authenticated); the local entry needs neitherghnor a network. - Both entries run on Windows. The one step that does not is escalating into the
e2e-failure-analyzerscripts, which shell out tounzip(references/evidence-escalation.md). Nothing in this skill needsunzip: local evidence reads traces withyauzlinstead. - CI entry only:
E2E_INSIGHTS_API_KEYset, or present in the repo-root.env.e2e(the query script falls back to it automatically). Get it from 1Password atop://Positron/E2E_dashboard_api_key/credential; without 1Password access, ask the Positron QA team.triage-history.jspre-flights this and exits withcause: "missing-api-key"plus the setup steps -- relay those steps and stop; this is not a triage finding and there is no degraded mode. Acause: "api-unreachable"is the different case: a key was found, so retry rather than sending the engineer to 1Password. - Neighbor skills:
e2e-failure-analyzer(itse2e-query-history.jsande2e-process-s3.jsare invoked directly); at the fix stagepositron-pr-helper,author-vitest-tests,author-e2e-tests.
Scripts
Run from the repo root. Flags and output contracts:
references/scripts.md. If a script itself breaks, see
references/script-fallbacks.md.
| Script | Use it to |
|---|---|
resolve-test-key.js |
turn a title / spec path / spec:line / dashboard URL into the exact test key |
triage-history.js |
get the failure patterns (dual-branch, merged, one occurrence each) |
find-prior-triage.js |
check whether this spec was triaged before |
fetch-pattern-evidence.js |
pull evidence for one occurrence of the selected pattern (CI) |
collect-local-evidence.js |
build the same summary from this machine's test-results/ (local) |
checkpoint.js |
start / resume / status; --set phase=X auto-derives nextAction |
record-diagnosis.js |
append the diagnosis block; it is what unblocks phase=done |
Start or resume
/debug-e2e-test "<test>" -- start the CI entry. <test> can be
anything that names one test: a leaf title, a spec path, spec.test.ts:41, a
full testName|||specPath key, or a dashboard URL.
/debug-e2e-test --local -- start the local entry (see below).
/debug-e2e-test with no argument -- don't ask which test: run the local
collector unfiltered. It ranks failures first and newest first, so it selects the
most recent local failure on its own, and names it back to you for confirmation.
Only when that finds nothing (no-results) is there a test name to ask for.
/debug-e2e-test --resume <triage-id> -- resume.
/debug-e2e-test --status -- list saved triages.
On --local (or when the engineer describes a failure they just produced):
skip straight to evidence -- no key resolution, no history, no checkpoint.
node .claude/skills/debug-e2e-test/scripts/collect-local-evidence.js
Act on its verdict, then read summaryFile and go to Determine root cause.
references/local-evidence.md owns the verdict
table, the run-it offer for no-results, what local evidence cannot answer, and
when to start checkpointing after all. The rest of this section is the CI entry.
On --resume: run
node .claude/skills/debug-e2e-test/scripts/checkpoint.js --triage-id <id> --read,
validate it, and continue from phase / nextAction. Do not repeat
completed history or evidence work unless the engineer asks to refresh, the
saved data is invalid, or the branch/test identity changed.
On --status:
node .claude/skills/debug-e2e-test/scripts/checkpoint.js --status.
Otherwise (new triage):
Resolve the test identity. Never hand-assemble the key -- pass whatever the engineer gave you (leaf title, spec path,
spec.test.ts:41from a stack trace, full key, or a pasted dashboard URL) straight through:node .claude/skills/debug-e2e-test/scripts/resolve-test-key.js '<whatever they gave>'It reads the real hierarchy from Playwright (~2s), so the
describenesting is never guessed. Useresolved.testKeywhenresolvedis non-null; when it's null, presentcandidatesand ask which -- do not pick one yourself. OninWorkingTree: falseor a non-nullnote, relay it. If it exits non-zero, readreferences/history-query.md.Run the history helper:
node .claude/skills/debug-e2e-test/scripts/triage-history.js \ --test-key '<testName>|||<specPath>' --lookback-days 14Act on its
verdict.stop: true(zero-runs-both,clean) or anerrorfield means stop and report -- readreferences/history-query.mdfor what each verdict means. Otherwise continue.Initialize a checkpoint and record the patterns.
<id>is thetriageIdfrom the history output, used verbatim in every later command -- inventing one silos the checkpoint from the work dir already holding the history and evidence, which is what--resumereads:node .claude/skills/debug-e2e-test/scripts/checkpoint.js --triage-id <id> \ --init --test-key '<key>'Check for prior triage before presenting the table:
node .claude/skills/debug-e2e-test/scripts/find-prior-triage.js \ --spec-path '<specPath>' --triage-id <id> \ --occurrence-shas '["<sha1>","<sha2>"]'A non-
noneverdict changes the plan -- readreferences/prior-triage.md. A merged fix additionally needs a denominator before "it held" is sayable: re-run the history query with--since-fix <mergedAt>and pass the resultingfixHeldnumbers back in. Without them the verdict istoo-recent-to-tellby construction, so skipping this step silently downgrades the answer.noneis not conclusive -- it matches spec paths, so a POM/helper-only fix never registers. If the failing locator is gone from the working tree, read that file anyway.open-attempt-in-flightmeans stop and point at the open PR.Present the failure modes as a table (never a run-on sentence). The Rate column comes from each pattern's
ratesarray, nevercount / totalRuns; Last seen renderslastSeenas5d ago (Jul 24), or justtodayat 0d. Include a "Seen on" column whenever two branches were queried:# Failure mode Count Rate Last seen Environments Seen on A locator.clicktimeout:.codicon-maximize4 100% on feature/x 5d ago (Jul 24) ubuntu/chromium feature/x only B toBeVisible()timeout:getByLabel('...')3 1.9% on main today ubuntu/electron main only Count and Rate are cumulative over the lookback, so they cannot separate an acute burst a merged fix already closed from an ongoing drip. A pattern whose
daysAgois stale next to the others is an already-fixed candidate: say so in your recommendation instead of steering the engineer there on count alone.When
lastSeen.dateisnullfor every pattern, render Last seen asunknownand give the recency read fromonsetinstead ("Started yesterday"). Report a dateless column as a gap in that one read, never as a reason the already-fixed question can't be asked.Ask which pattern to prioritize whenever the table has more than one row. Give your own read ("A is dominant at 99% -- start there, or focus on B?") but let the engineer decide; they may already know which failure they care about. A single pattern needs no choice. Save the selection to the checkpoint (
--set selectedPattern=A --set phase=pattern-selected).
Investigate the selected pattern
Local entry: collect-local-evidence.js already wrote your summary.md --
start at step 2. Steps 2 and 3 are shared.
- Fetch evidence for the pattern's representative occurrence:
(The helper strips thenode .claude/skills/debug-e2e-test/scripts/fetch-pattern-evidence.js \ --report-url '<representativeOccurrence.report_url>' \ --triage-id <id> --pattern Aindex.html#?testId=fragment and filters the report to this one test itself.) - Read the generated
summary.md(failure, timeline tail, sibling tests, error-shaped logs, unresolved questions). Read only the summary first. - State the concrete questions that remain. Before each escalation past the
summary, show the evidence block (
Question/Next artifact/Reason) defined inreferences/evidence-escalation.md-- can't fill all three fields, don't escalate. Show it, then dispatch the escalation as a subagent with that block as its prompt; don't open the artifact yourself, and don't let the block reach only the subagent. That reference owns the block format, the dispatch contract, the ladder, the reasons a second occurrence is allowed, raw-log spelunking, and 403/null handling. - Save
phase=evidence-gatheredto the checkpoint.
Determine root cause
This is a collaborative dig, not a rubber-stamped verdict. Read
references/triage-rubric.md -- the taxonomy,
the dismissal bar, what each evidence type proves, and the locator-drift
decision.
State: the observed mechanism (citing trace step / log line / snapshot); what the evidence rules in and out; the surviving alternatives; and a fix that could plausibly change the failure rate (a fix that couldn't is not a fix -- keep digging).
Delegate cross-file tracing to an Explore subagent only after the evidence
names a concrete symbol / selector / event / subsystem -- under the same cap and
forbidden list as an evidence read ("Delegate the read" in
references/evidence-escalation.md).
Then agree the fix approach, before the first edit. Table the plausible fixes
(approach / what it changes / risk) with your recommendation and let the
engineer pick -- editing before the pick burns a context on a rejected
approach. Carry the pick as diagnosis.fixApproach.
Save the diagnosis to the checkpoint (--patch a diagnosis object) and set
phase=hypothesis-ready. Include the fields record-diagnosis.js renders
(confidence, summary, targetedFailure, signal, hypothesis, optional
supersedes) plus fixApproach from the gate above -- see
references/diagnosis-block.md for what each
must contain.
Reproduce and fix
Checkpoint the diagnosis before implementing, then set
phase=implementation. That is the required step: history, evidence, and
working-tree edits are all durable on disk, so the invariant to hold is that
implementation could start from the checkpoint alone -- a fixApproach
specific enough to act on without re-reading the evidence.
That invariant is what makes a clear safe, so you never have to propose one: the
engineer clears when they want and --resume <id> picks the triage back up.
Read references/reproduction.md now -- it owns
keeping this phase's context small, project choice, race verification, and the
RED bar
you must hold when author-vitest-tests writes a lower-level regression
test (that skill drives toward green; it does not enforce RED-first).
Record the result and close out
Every triage ends by declaring an outcome and recording its diagnosis -- this
is not optional, and checkpoint.js refuses phase=done until it's satisfied.
On the local entry there is no checkpoint to gate you, so the rule is yours
to keep: a PR or an issue still gets the block. A local dig that ends with a fix
and no artifact ends when the fix is verified -- say so and stop; don't
manufacture a checkpoint to close.
The outcome spans two axes (what you found x what you did):
| Outcome | Meaning | Where the block goes | To reach done |
|---|---|---|---|
fix-test |
test bug, fixed in a PR | the PR | record-diagnosis.js --pr <n> --outcome fix-test |
fix-product |
product bug, fixed in a PR | the PR | record-diagnosis.js --pr <n> --outcome fix-product |
file-issue |
product bug, filed not fixed | the new issue | record-diagnosis.js --issue <n> --outcome file-issue |
no-op |
not fixed and not filed (accepted flake, dup, backlog, handed off) | checkpoint only | --set outcome=no-op --set outcomeReason="..." |
outcome is the primary artifact -- a secondary note (e.g. mentioning a
product race in the backlog while you fix the test) does not change it. When a
triage genuinely produces two artifacts, the block goes on both and outcome
still names one: references/diagnosis-block.md.
A returning sub-tool is not the end of the triage -- opening the PR via
positron-pr-helper or a passing author-vitest-tests run resolves a step.
Once the PR/issue exists:
record-diagnosis.js --triage-id <id> --pr <n> --outcome <fix-test|fix-product>(or--issue <n> --outcome file-issue) appends the block and setsoutcome+outcomeRef+diagnosisBlockRecordedin one call. For ano-op, skip this andcheckpoint.js --set outcome=no-op --set outcomeReason="..."instead.checkpoint.js --set phase=done.