Start your first response with the 🧪 emoji.
Absolute Deflake
Find tests that pass and fail nondeterministically, diagnose the root cause of each,
and fix it — not by retrying or skipping, but by removing the source of nondeterminism.
Output is evidence (failure rate per test) → cause → fix, verified by repeated runs.
Runs the shared engine in references/health-engine.md — read it for the
DETECT → SCAN → TRIAGE → FIX → VERIFY → REPORT loop and the safety contract. This file
covers only what's specific to flaky tests.
When to use
- "Our CI is flaky", "this test fails randomly", "fix the intermittent failures".
- A test passes locally but fails in CI (or vice versa), or fails ~1 in N runs.
- Burning down a backlog of
retry/skip-marked tests that mask real flakiness.
Not for tests that fail deterministically — that's a real bug or a real regression
(/absolute work for a fix, or just fix it). deflake targets nondeterministic failures.
What it scans
Establish flakiness empirically — a test isn't flaky because someone said so. Use
preferences.health.deflakeRuns from config as the default N for repeat-runs (else 20):
| Ecosystem |
Repeat-run / detect |
| Jest/Vitest |
run suite N× (--run loop), randomize order (--shuffle / testSequencer) |
| pytest |
pytest-randomly + pytest --count=N (pytest-repeat); -p no:randomly to A/B |
| Go |
go test -count=N -shuffle=on ./..., -race |
Also mine signals: existing retry/flaky/skip annotations, CI history if reachable, and
run the suite both in isolation and in full/parallel — order- and concurrency-
dependent failures only show one way. Record a failure rate per suspect test.
Common root causes (diagnose, don't guess)
| Cause |
Tell |
Fix |
| Test-order / shared state |
passes alone, fails in suite (or vice versa) |
isolate state; reset/teardown between tests |
| Time / clock |
fails near midnight, DST, or under load |
fake timers / inject clock; no real sleep |
| Async race / missing await |
fails under parallelism or slow CI |
await the actual condition; no fixed timeouts |
| Randomness |
fails ~X% with no pattern |
seed the RNG; fix the seed in tests |
| Network / external I/O |
fails offline or on slow links |
mock/stub the boundary |
| Unordered collections |
fails on map/set iteration order |
sort before asserting |
| Resource leak / port reuse |
fails on repeat or parallel runs |
unique resources; clean up |
Risk ranking (TRIAGE)
| Wave |
Class |
Default |
| 1 |
clear, isolated cause (seed, await, fake clock, sort) |
fix now |
| 2 |
shared-state / ordering — needs fixture refactor |
fix this pass, per test |
| 3 |
flakiness pointing at a real product race, not just the test |
gated — surface; may be a genuine bug to fix in code |
A flaky test sometimes means the code has a race, not the test. Don't "stabilize" the test
into hiding a real concurrency bug — flag wave-3 cases for a real fix.
Fix & verify
- Fix the cause. Then prove it: re-run the test many times (and shuffled / parallel /
with
-race) — green once is not deflaked; green across N randomized runs is.
- Remove the
retry/skip/flaky annotation that was masking it once the cause is fixed.
- Never "fix" by adding retries, raising timeouts blindly,
sleep, or skipping the test —
that hides flakiness, doesn't remove it.
- Re-run the full suite to confirm the fix didn't destabilize neighbors.
Gotchas
- Retry/skip as a fix. Masks the flake, ships the nondeterminism. Forbidden here.
sleep to dodge a race. Slows the suite and still flakes under load. Await the condition.
- One green run = done. Flakes are probabilistic — verify with many randomized runs.
- Stabilizing a real product race. If the code races, fix the code, not just the assertion.
- Ignoring order/parallel dimension. Run isolated and in-suite; the bug hides in whichever you skip.
Companion commands
/absolute upgrade — a flaky suite makes upgrade verification unreliable; deflake first.
/absolute debt — flaky-test annotations are test debt; this clears them at the root.
/absolute work — when the flake is a genuine product-code race needing real design.
1---2name: absolute-deflake3description: Flaky test fixes: detect nondeterministic tests empirically (repeat/shuffle/parallel runs), diagnose the root cause, fix it — never retry/skip/sleep — and verify across many randomized runs. Triggers on "absolute deflake", "fix flaky tests", "CI is flaky", "this test fails randomly/intermittently".4license: MIT5---67> Start your first response with the 🧪 emoji.89## Absolute Deflake1011Find tests that pass and fail nondeterministically, diagnose the **root cause** of each,12and fix it — not by retrying or skipping, but by removing the source of nondeterminism.13Output is evidence (failure rate per test) → cause → fix, verified by repeated runs.1415Runs the shared engine in **`references/health-engine.md`** — read it for the16DETECT → SCAN → TRIAGE → FIX → VERIFY → REPORT loop and the safety contract. This file17covers only what's specific to flaky tests.1819---2021## When to use2223- "Our CI is flaky", "this test fails randomly", "fix the intermittent failures".24- A test passes locally but fails in CI (or vice versa), or fails ~1 in N runs.25- Burning down a backlog of `retry`/`skip`-marked tests that mask real flakiness.2627Not for tests that fail *deterministically* — that's a real bug or a real regression28(`/absolute work` for a fix, or just fix it). `deflake` targets *nondeterministic* failures.2930---3132## What it scans3334Establish flakiness empirically — a test isn't flaky because someone said so. Use35`preferences.health.deflakeRuns` from config as the default N for repeat-runs (else 20):3637| Ecosystem | Repeat-run / detect |38|---|---|39| Jest/Vitest | run suite N× (`--run` loop), randomize order (`--shuffle` / `testSequencer`) |40| pytest | `pytest-randomly` + `pytest --count=N` (`pytest-repeat`); `-p no:randomly` to A/B |41| Go | `go test -count=N -shuffle=on ./...`, `-race` |4243Also mine signals: existing `retry`/`flaky`/`skip` annotations, CI history if reachable, and44run the suite both **in isolation** and **in full/parallel** — order- and concurrency-45dependent failures only show one way. Record a **failure rate** per suspect test.4647---4849## Common root causes (diagnose, don't guess)5051| Cause | Tell | Fix |52|---|---|---|53| Test-order / shared state | passes alone, fails in suite (or vice versa) | isolate state; reset/teardown between tests |54| Time / clock | fails near midnight, DST, or under load | fake timers / inject clock; no real `sleep` |55| Async race / missing await | fails under parallelism or slow CI | await the actual condition; no fixed timeouts |56| Randomness | fails ~X% with no pattern | seed the RNG; fix the seed in tests |57| Network / external I/O | fails offline or on slow links | mock/stub the boundary |58| Unordered collections | fails on map/set iteration order | sort before asserting |59| Resource leak / port reuse | fails on repeat or parallel runs | unique resources; clean up |6061---6263## Risk ranking (TRIAGE)6465| Wave | Class | Default |66|---|---|---|67| 1 | clear, isolated cause (seed, await, fake clock, sort) | fix now |68| 2 | shared-state / ordering — needs fixture refactor | fix this pass, per test |69| 3 | flakiness pointing at a **real product race**, not just the test | gated — surface; may be a genuine bug to fix in code |7071A flaky test sometimes means the *code* has a race, not the test. Don't "stabilize" the test72into hiding a real concurrency bug — flag wave-3 cases for a real fix.7374---7576## Fix & verify7778- Fix the **cause**. Then prove it: re-run the test **many times** (and shuffled / parallel /79 with `-race`) — green once is not deflaked; green across N randomized runs is.80- Remove the `retry`/`skip`/`flaky` annotation that was masking it once the cause is fixed.81- **Never** "fix" by adding retries, raising timeouts blindly, `sleep`, or skipping the test —82 that hides flakiness, doesn't remove it.83- Re-run the **full** suite to confirm the fix didn't destabilize neighbors.8485---8687## Gotchas88891. **Retry/skip as a fix.** Masks the flake, ships the nondeterminism. Forbidden here.902. **`sleep` to dodge a race.** Slows the suite and still flakes under load. Await the condition.913. **One green run = done.** Flakes are probabilistic — verify with many randomized runs.924. **Stabilizing a real product race.** If the *code* races, fix the code, not just the assertion.935. **Ignoring order/parallel dimension.** Run isolated *and* in-suite; the bug hides in whichever you skip.9495---9697## Companion commands9899- **`/absolute upgrade`** — a flaky suite makes upgrade verification unreliable; deflake first.100- **`/absolute debt`** — flaky-test annotations are test debt; this clears them at the root.101- **`/absolute work`** — when the flake is a genuine product-code race needing real design.