There is a checkout test in a suite that fails about one run in four, and has done for months. It fails on CI, passes on the retry, and everyone has learned the ritual: the build goes red, somebody hits re-run, the second attempt is green, the pull request merges. The test has been called flaky so many times that the word has stopped meaning anything. Last week an order was charged twice in production — a race the test had been reporting, in the only way a failing test can report something, for months.
The test was never flaky in any useful sense. It was intermittently correct, which is the most valuable thing a test can be and the exact category teams are trained to stop reading. This note is about the difference between those two descriptions, and about the two mechanisms that decide which one your team ends up believing: retries, and quarantine.
A flake is a pass and a fail on the same code
Start with a definition that removes the ambiguity. A test is flaky when it produces different results against identical code — same commit, same source, both a pass and a fail in the record. That is not an opinion about the test; it is a fact you can measure, and it means the suite is being decided by something other than the thing under test.
There are exactly two places that something can live, and telling them apart is the whole job:
- The nondeterminism is in the test or the environment. A snapshot read where an assertion was needed, data two parallel workers share, an unmocked third-party script, a clock that crosses midnight, a runner under memory pressure. The product is fine; the measurement is not.
- The nondeterminism is in the product. A genuine race — a double submit, an optimistic UI that reports success the server never confirmed, a read that can happen before the write it depends on. The test is doing its job, badly, and the defect ships.
A green retry cannot distinguish those two. It does not know which one it just suppressed, and it never will, because the evidence it would need was the failing run — the trace, the network log, the DOM snapshot at the moment of failure — and the retry overwrites it with a passing one.
A retry answers one question: can this test ever pass? It never answers the question you need: is the thing underneath it broken?
Retries convert a finding into a green checkmark
Retries are not useless; they are a tool that got promoted into a policy. Set to absorb genuine infrastructure blips — a runner that lost its network for two hundred milliseconds — they are reasonable. Set globally and left there, they become the mechanism by which a suite reports success while measuring something broken.
The arithmetic is unforgiving. Take a test that fails one run in four. With two retries, three attempts, the build goes red only when all three attempts fail, which is about 1.6% of runs. The suite reports green on 98% of runs of a defect that a quarter of your users would hit on their first attempt. Nothing has been stabilised. The failure has been moved to a place the pipeline does not look.
Two habits keep retries honest rather than corrosive. Cap them, and make the cap explicit, because an unbounded retry is a suite that cannot fail. And keep a retry-passed test as its own outcome rather than folding it into "passed". Playwright already draws this line for you: a test that fails and then passes on retry is reported as flaky, which is neither a pass nor a fail, and a report you can filter on. The failure is not in the runner. It is in the pipeline that reads a flaky count of zero and prints a green tick.
So the first change is not a threshold or a tool. It is that a flake becomes a named work item, and the question asked about it is which of the two categories above it belongs to — not whether re-running it works. It does. That is the problem.
Classify first, then fix
Trying to fix a flake without classifying it is how teams add a wait, remove a wait, add it back, and eventually quarantine the test with no more understanding than when they started. There are a small number of root causes, and each leaves a different signature in the record. Read the signature first — the trace of the failing attempt, not the passing one.
| Signature | What it usually is | Where the fix belongs |
|---|---|---|
| Intermittent timeout on an action; the element was present in the trace | The assertion read a snapshot instead of polling, or an async call was never awaited | Test: web-first assertions, and lint for floating promises |
| Passes alone, fails in the full suite or at higher parallelism | Shared state — a user, a row, a file, a stub server two workers both own | Test: per-test data, a real reset boundary, isolated authentication |
| Fails only on CI, never locally, worse under load | Runner contention. The test is not slow; the machine is | Infrastructure and budget: fewer workers per runner, sane timeouts |
| Fails on a schedule: midnight, month end, a timezone boundary | The test is reading the real clock or the real locale | Test: freeze time, pin timezone and locale |
| Fails when a third party is slow or rate-limits | An unmocked external dependency donating its uptime to your pipeline | Stub at the network layer, or quarantine it and label it as external |
| Fails on a double submit, a duplicate row, a lost update | Not a test problem. A product race the test is correctly catching | Product: this is the one you wanted to find |
Two rows are the ones teams get wrong for a long time. The last one is a defect and belongs on the bug board, not in the test backlog. The environment row is more insidious, because "passes locally, fails on CI" gets diagnosed as flakiness by default — and a test failing because the runner ran out of memory is not a test problem at all. Making the suite lighter or the runner bigger is a legitimate fix, and treating it as a test issue spends a week in the wrong repository.
What auto-waiting covers, and what it does not
Modern browser runners removed an entire category of flake by waiting before acting. Before a click, Playwright checks that the element is attached to the DOM, visible, stable across frames, enabled, and actually receiving pointer events. None of that needs a sleep, and the old habit of pausing for a fixed two seconds is a guess that is sometimes too short and always too long.
The boundary is sharp, and most remaining timing flake sits just past it. Auto-waiting understands actionability. It has no idea about your application's semantics — that a dashboard painted skeleton rows before real ones, that a list re-sorts after it renders, that a button enables itself before the request that validates the input has returned. In those cases the element is genuinely ready to be clicked and the application is not ready to be clicked. The fix is to assert the outcome a user would see, and let the assertion poll until reality matches, rather than asserting on whichever intermediate state happens to exist right now.
That distinction explains the flake that is never fixed by adding a wait:
// Reads the DOM once, at this instant, and compares. If the UI has not
// updated yet the assertion fails, even though it would pass 100ms later.
const status = await page.getByTestId('status').textContent();
expect(status).toBe('Paid');
// Polls the same locator until the condition holds or the timeout expires.
await expect(page.getByTestId('status')).toHaveText('Paid');
Assertions that poll are the difference, and the pattern extends past element state. locator.count() and locator.all() return a snapshot rather than a live view, so a loop over a list that is still loading iterates the wrong number of items — wait for the expected count first, then iterate. Where a value settles over time rather than a single element's state, retry the whole check instead of the individual read.
Register the listener before the action
The other timing defect is order, not duration. Waiting for a response after triggering the request is a race you lose whenever the server is fast:
// The response can arrive before anything is listening for it,
// and the wait never resolves.
await page.getByRole('button', { name: 'Place order' }).click();
const response = await page.waitForResponse((r) => r.url().includes('/api/orders'));
// Listen first, then act. The promise captures the response whatever the timing.
const responsePromise = page.waitForResponse((r) => r.url().includes('/api/orders'));
await page.getByRole('button', { name: 'Place order' }).click();
const response = await responsePromise;
expect(response.status()).toBe(201);
The same rule covers stubbing: register the route handler before the navigation that triggers the request it is meant to serve. This class of bug is invisible on a slow local backend and appears on fast CI runners — exactly backwards from the intuition that a faster environment should be more reliable.
Isolation is a property of the suite, not of the test
The shared-state flake is the one that survives every timing fix, because it is not a timing problem. The diagnostic is specific and worth memorising: run the suspect test alone and it passes; run the suite and it fails. That is not a slow test. That is a test whose inputs were changed by another test — a user whose data the neighbouring worker just mutated, records from a previous run still sitting in the table, an assertion that assumes "the newest item is mine".
Parallelism makes it worse in a straight line: the more workers you add to keep the suite fast, the more collisions between tests that each assumed they were alone. Three habits remove most of it.
- Every test creates what it needs. Not "a user exists", but this test's user, named after the test, carrying the data the assertions expect. A shared seed fixture is the most common source of order dependence in a browser suite.
- Reset at a real boundary. Before each test, reset through the API or a transaction, not through the UI. UI setup makes the test's start state depend on the application being healthy, so when the app is broken the failure lands in setup and reads like an infrastructure error.
- Reuse authentication, not state. Load a pre-authenticated session rather than logging in through the form on every test. The login flow is slow, identical for every test, and not what any of them are testing.
There is a temptation to see an order-dependent failure as proof of a concurrency defect in the product, and occasionally that is exactly right — but only after you have ruled out the suite sharing the database row, the file, or the stub. Isolation first, then ask whether the collision is in your tests or in your transaction handling.
Remove the inputs you did not choose
Two more sources of nondeterminism have nothing to do with timing and are cheaper to remove than to investigate, so remove them before they cost you a day of hunting something that will not reproduce.
Time. Any assertion that depends on the current date, the day of the week, or the local timezone will eventually fail — once, on a month boundary, at midnight, over a leap day — which is precisely the shape of a flake that never reproduces. Freeze the clock in the browser and drive it forward deliberately: a five-minute session timeout becomes a test that runs in milliseconds and fails on time rather than on luck. Pin the timezone and locale of the run as well, so a machine in another region cannot change the answer.
Randomness and version drift. Seed every random source a test depends on. Pin the browser and runner versions, because a suite that runs on whatever image happens to be current this month is a suite whose results cannot be compared across weeks — and an engine version change produces exactly the "it passed yesterday" signature that gets filed as flake.
One policy does more than any of that: treat a fixed sleep as a banned pattern, not a discouraged one. A rule that says "avoid sleeps" leaves room for the case where this one is genuinely needed, and the case is almost never the case. A ban removes the argument outright, and a lint rule that fails on un-awaited async calls removes the other half before it can be committed.
Quarantine is not a quieter retry
You now have a test you cannot fix today and cannot let hold the merge queue. The instinct is a bigger retry count. The right mechanism is a quarantine lane, and the difference between the two is the only thing that matters here: a retry hides the evidence; a quarantine files it.
Take the test out of the blocking suite and put it somewhere it still runs on every commit, still reports, and does not gate the merge. It keeps producing the failing run — the trace, the classification, the frequency — which is the data you need to fix it. A retry produces the opposite: a green run that replaces the red one.
Four properties separate a quarantine that works from one that becomes a graveyard, and every failed quarantine process is missing at least one of them.
- The quarantined test still runs. A skipped test is deleted with extra ceremony and a nicer feeling. It must execute on every run and its result must be recorded, or the coverage has been thrown away without anyone admitting it.
- Entry comes from history, never from one red run. A test that failed once had a bad day. A test with both a pass and a fail recorded against the same commit is flaky by definition, and that is a machine's judgement rather than a tired engineer's.
- Every entry has an owner, a ticket and a date. Enforce it in CI — an entry added without a ticket reference should fail a check of its own. A quarantine with no clock is a retry wearing a different name; the date is what turns "we will get to it" into somebody's work item.
- Exit requires evidence. A test returns to the blocking lane when it has passed a deliberate run of consecutive attempts, not when it happens to be green today.
The lane itself is ordinary pipeline configuration — two projects, one gating and one reporting, with retries disabled in both so a failure in either lane is honest:
// playwright.config.ts
export default defineConfig({
projects: [
{
name: 'blocking', // Gates the merge. A red run stops the pull request.
testIgnore: '**/*.quarantine.spec.ts',
retries: 0, // A retry here would recreate what we are removing.
},
{
name: 'quarantine', // Runs and reports on every commit. Does not gate.
testMatch: '**/*.quarantine.spec.ts',
retries: 0,
},
],
});
Turning retries to zero inside the quarantine lane is not an oversight. A retry there would smooth over the exact failure the test was moved to study, one level down. Let it be red, and read the pattern.
Two guardrails stop the lane becoming a permanent hole in the suite. Track the number of quarantined tests as a reliability metric in its own right, and treat growth as a signal — flakiness is outrunning the team's ability to fix it, which is a capacity fact worth acting on rather than a testing detail. And cap the lane, weighted by what the tests protect: five quarantined checkout tests are a bigger blind spot than forty tooltip assertions, and a flat count cannot tell the difference. When the weighted cap is hit, stop admitting entries until the backlog shrinks.
Deletion is a legitimate exit, and saying so out loud is what keeps the process honest. A test that duplicates coverage one level down, or guards something nobody would notice breaking, is worth less than the runtime and the attention it costs. Removing it deliberately is a decision. Leaving it quarantined forever is also a decision — just one nobody made.
Defend the gate, or you have not built one
Every mechanism above rests on the blocking lane being real, and the one configuration error that silently reverses all of it is a branch-protection rule that does not list the blocking lane as required. A gate that is not required is a report.
So verify the separation the way you would verify anything with a consequence: open two deliberately broken pull requests. Break a test in the blocking lane and confirm the merge is refused. Break one in the quarantine lane and confirm the merge is allowed. That pair of runs is the contract, it takes ten minutes, and it is worth repeating whenever branch protection changes — because a single mislisted check is invisible until the day it matters.
The end state is a suite trusted at zero retries. Not because retries are evil, but because a suite you have to retry is a suite whose red and green both mean less than they should, and once engineers learn that red sometimes means nothing, they stop reading. That is the failure that actually costs you: not the flaky test, but the pipeline nobody believes. We described the same collapse from the other direction in our note on eval gates that lose their authority, and it is the same mechanism — a check that reports unreliably is worse than no check, because it charges the upkeep and returns false confidence.
The suite has to stay fast enough to be run
One structural point sits under all of the above. A browser suite is the most expensive way to check a behaviour, and its cost is runtime — so a suite that takes forty minutes gets skipped, or sharded until nobody waits for it, and a skipped gate is not a gate. The end-to-end layer earns its place on the flows where being wrong costs money: checkout, payments, authentication, the one mutation that cannot be taken back. Everything below that belongs in unit and integration tests that run in seconds and fail deterministically.
That is the same rule we apply to coverage in client work — aim it at risk rather than at a percentage, with a small set of end-to-end tests over the flows that would hurt and the volume of checking pushed down a level where it is cheap and exact. A browser suite budget and a flake programme are the same budget: every test you move out of the browser is a flake you will never have to classify.
What this means in practice
Order the work by what the evidence supports, not by what is easiest. Stop folding retry-passed runs into the pass count, and record a flake as its own outcome with a name attached. For each one, read the signature before touching anything: timing, isolation, environment, a clock, or a genuine race in the product. Fix the first four where they live, and move the fifth to the bug board — the test has already done its job.
Then remove the inputs you did not choose — freeze time, pin versions, seed randomness, ban the fixed sleep — and isolate the data so every test owns what it asserts on. Take whatever remains off the merge gate into a lane that still runs, with an owner, a ticket, a deadline, and evidence required to come back. And prove the gate with two deliberately broken pull requests, because the separation between blocking and reporting is what all of it depends on.
None of this is exotic, and none of it is about the model or the framework. It is the ordinary discipline of a measurement you intend to act on: know what varies, control it where you can, and keep the failures you cannot yet explain somewhere a person is still looking at them. If your suite has a test everyone re-runs without reading, tell us what it is doing — cleaning that up is routine work on the delivery pipelines we build, and it starts in the same place as our platform and systems engineering. Two related notes: generated tests and the coverage that catches nothing covers the same trust problem from the unit-test side, and reviewing AI-generated code is where the banned-sleep policy and the isolation rules get enforced, because a flake nobody notices at review time is a flake somebody else spends a day on later.