Technology Build Note

When Your Own Tests Are the Thing That's Wrong

The application was fine. One CI failure was real, three later diagnoses were wrong, and the test was asking for state before it had loaded.

A test asserting synchronously on state that loads asynchronously, and the awaited assertion that fixes it

Green here, red there

I moved HGV HUB’s test suite onto independent CI. Locally it was green, every time. On the runner it was red.

Not catastrophically red — 384 of 386 passing on a representative run, two Vehicle Check flows failing on “Unable to find an element”. The tempting reading is that CI is fussy.

The detail that ruled that out: the failing set kept changing between runs. Sometimes those two flows, sometimes another as well; once a diagnostic run passed outright. A consistent environment difference does not behave like that. Something was racing, and my machine was winning often enough that I had never seen it.

The first diagnosis was right — about something else

The first CI failure had nothing to do with any of that.

The workflow pinned Node 20. The test environment pulls in a version of undici that calls an API newer than Node 20 provides, so the suite fell over before reaching an assertion. I develop on Node 24, which is why I had never seen it.

I moved CI to Node 24 and declared the engine requirement so the two environments would agree rather than discovering the mismatch through a red build.

That was correct. It fixed a genuine compatibility problem, and no test was weakened to make it pass. Worth being clear about, because it is easy to mislabel in hindsight: the first diagnosis was not a bad guess. It solved a real failure and revealed a different one underneath.

Then three plausible explanations

Each of the next three was reasonable. All three were wrong.

The runner is slower. Measurably true — the suite takes roughly two and a half times longer there. The two failing flows are the heaviest and ran close to Testing Library’s default async timeout locally, a margin that plausibly does not survive slower hardware. I raised that timeout fivefold. No effect.

State is leaking between tests. Also plausible: these flows use IndexedDB, and stale state is a classic cause of order-dependent failures. I added a stronger teardown barrier. It made things worse — the extra transactions pushed the slowest file from two failures to three. Rejected, and reverted.

It is the test runner’s timeout, not the query’s. I was confident enough about this one to title the commit “Root cause: vitest testTimeout, not state leakage”.

It was not. A controlled run that stripped everything else out and kept only that change still failed, eleven minutes later.

I quote my own commit subject because it is the useful part. Nothing about it was careless — it named a mechanism, excluded an alternative, and was written the way you write something established. Confident wording is not evidence. The only thing separating it from a real finding was an experiment I had not run.

Stop reasoning, start measuring

At that point I had produced four explanations and fixed one problem. That was the pattern: generating stories about the runner from a machine where the failure would not occur.

So I stopped reasoning and started buying evidence. A throwaway diagnostics branch, pushed repeatedly, instrumenting one flow at a time — ten CI runs in twelve minutes, inside the environment that was actually failing.

That is what changed the outcome, and the number matters less than the shift: each run answered one narrow question rather than confirming a narrative. The decisive one removed the diagnostics and kept only the supposed timeout fix. It failed — killing the hypothesis cleanly and leaving the assertion as the only place I had not properly looked.

getBy does not wait

Both failing flows did the same thing:

await screen.findByRole("heading", { name: "New Vehicle Check" });
expect(screen.getByText("Tyres")).toBeInTheDocument();

The heading renders as soon as the page mounts. The checklist items do not — they arrive from an asynchronous IndexedDB read a tick later.

findByRole waits. getByText checks immediately. So the test quietly assumed: because one awaited element is on screen, all the async state this flow depends on must have arrived too.

LOCAL heading data assert PASS CI heading assert data FAIL FIXED heading await data PASS

The assertion runs at the same point every time. The only thing that changes is whether the data got there first — and whether the query is willing to wait for it.

On a fast machine that read had usually settled by the time the heading appeared, so the assertion passed on timing luck rather than correctness. On a slower runner it had not. The assertion always ran at the same logical point; what varied was whether the data got there first.

Which explains why two timeout changes moved the symptom without curing it. No timeout can help a query that never waits. The async timeout only applies to queries that wait; the runner timeout only changed which error I saw.

What I actually changed

Two assertions. getByText became await findByText. Same text, same required element, same end state — the assertion did not get weaker, it got correctly synchronised with what it was asserting.

No application code changed at all. The checklist loaded correctly the whole time. The test was asking for it too early.

Both timeout changes stayed, and the reason matters, because the lesson is not “timeouts are bad”. The async timeout is defensible on its own merits; the runner timeout is real headroom for a suite that takes two and a half times longer there. Both were introduced for a wrong reason and kept for a right one. What was wrong was reaching for them instead of understanding what the assertion waited for.

The fix held: five consecutive green runs on the diagnostics branch, then green on the release — six in the environment that had exposed the problem, which is the only reason the count is worth mentioning.

Two habits came out of it. Before changing the environment again, ask what assumption the assertion is making. And when reasoning keeps producing plausible answers, make the next experiment cheaper than the next theory.

testingcidebuggingreact

More Notes like this


All Notes