Corrected 2026-08-09 — the mechanism first filed here was wrong, and the real one is a vacuous assertion rather than a missing one. Original text kept below the line for the record.
Actual mechanism
The spec DOES await the persist, at tools/simulation/ui/specs/i18n.spec.ts:95-100:
const persisted = page.waitForResponse((response) =>
response.url().includes("/api/v1/me/language")
&& response.request().method() === "PUT",
);
await selector.selectOption(lang);
expect((await persisted).ok(), `the ${lang} preference was not persisted`).toBe(true);
That assertion passed in the failing run — the failure came later, at line 156. So "the reload beat an un-awaited PUT" is refuted by the code.
What actually happens is that the listener matches any PUT /me/language, and there is another one in flight:
- Line 91 does
selector.selectOption("en") to normalise the starting language. That fires PUT {language:"en"} and is not awaited.
- Line 92 asserts the English nav is visible — instant and local, so it does not wait for that PUT.
- Line 95 registers
waitForResponse for any PUT to /me/language.
- Line 99 calls
selectOption(lang). LanguageSelector.tsx:44 renders <select … disabled={busy}>, so Playwright's actionability check waits for the en flight to finish. During that wait the en response arrives — and satisfies the listener registered at step 3.
- Line 100 therefore asserts
.ok() on the reset's response. It passes vacuously, while the lang PUT has only just been fired.
localStorage.clear() + reload() at 134-135 can now outrun the lang PUT, so /me returns no preference and the shell renders English — the observed failure.
es is test #1 and pays cold start, which is what makes the en PUT slow enough to still be in flight at step 3. On a warm run it has already settled, the listener correctly matches the lang PUT, and everything holds — hence one failure in twenty runs.
This is the #428 shape: a listener bound to the wrong request.
Fix
Bind the wait to the specific language, so the reset's response cannot satisfy it:
const persisted = page.waitForResponse((response) =>
response.url().includes("/api/v1/me/language")
&& response.request().method() === "PUT"
&& response.request().postDataJSON()?.language === lang,
);
putMeLanguage sends { language } (web/src/api/cluckwork.ts:420), so the payload is the discriminator. Ordering needs no extra handling: the select is inert during the en flight, so that PUT is always complete before lang is selected.
The vacuity is the real defect here — the assertion on line 100 exists precisely to prove the preference was persisted, and it can pass without ever observing that request.
Original report (mechanism since refuted)
Seen once on CI, on a docs-only PR whose diff cannot reach this job (#484 touches AGENTS.md and a comment in deploy/docker-compose.yml; the sim harness runs the self-contained tools/simulation/docker-compose.sim.yml). Same app bytes as main, which passed an hour later — so the discriminator is timing, not code.
Evidence
Failing job: run 31345832064 — 29 passed, 1 failed.
Passing job on main, same code: run 31348597389.
i18n.spec.ts:69 › switching to es renders that language across the shell
Error: the es preference did not survive a reload with the device hint cleared
— it was never persisted server-side
Locator: getByRole('link', { name: 'Informes', exact: true })
Timeout: 10000ms — element(s) not found
Durations for the same parameterized test:
| Run |
es (test #1) |
tl (test #2) |
Result |
main |
2.3s |
1.7s |
pass |
| PR #484 |
12.2s |
2.1s |
fail |
On main the whole test — switch, clear storage, reload, re-assert — finished in 2.3s, nowhere near the timeout. On the failing run the post-reload assertion burned the full 10s.
The uploaded screenshot shows the post-reload page is healthy and entirely English: English nav, and the Language selector itself reading English. Not a broken render, not a missing catalog key, not a crash — the app correctly showing English because no preference was persisted to find. The selector is enabled, so the flight had settled by then.
Permissions are ruled out: tl ran the same body with the same Read-only persona in the same run and passed.
Mechanism
web/src/session/LanguageSelector.tsx:33-38 fires both calls and awaits neither:
void i18n.changeLanguage(lang); // local, instant
void run("language", () => putMeLanguage(lang)).catch(…) // server persist, not awaited
The local switch repaints immediately, so the test's first assertions (Informes visible, <html lang="es">) pass off device state alone. The test then does, at tools/simulation/ui/specs/i18n.spec.ts:134-135:
await page.evaluate(() => window.localStorage.clear());
await page.reload();
with nothing awaiting the in-flight PUT /me/language. If the reload beats the PUT, nothing was persisted, /me returns no language, and the shell renders English — the observed failure. es is test #1 and pays cold start (service worker install, cold container), which is why it is the one that loses.
What is NOT proven
The PUT was never directly observed. The container-log dump is failure-only and tail-truncated (the failing run's dump covers only the last 19s, starting at 01:01:00, while the test ran ~01:00:21-33; the passing run dumped nothing at all), and the artifact upload carries only video.webm and the screenshot — trace.zip is generated but not uploaded, and the trace is what holds the network log.
So "the reload beat the PUT" vs "the PUT was issued and failed" is inferred from source plus timing, not observed. Same defect class and same fix either way, but the distinction is currently unfalsifiable on CI. See #NEXT for the artifact gap that made it so.
Proposed fix
LanguageSelector.tsx:44 already renders <select … disabled={busy}> — the control is inert while the persist is in flight, deliberately (see the comment on line 32). So the test has a deterministic signal: wait for the select to re-enable before clearing storage and reloading. Does not weaken the assertion, and matches the established pattern for this shape (wait on the disabled-while-pending control rather than on a fixed delay).
Related
Same family as #470 (arrange-phase probe that could not distinguish "not started" from "finished") and #428 (listener registered after the request it waits for): a test asserting state that depends on an un-awaited async write.
Corrected 2026-08-09 — the mechanism first filed here was wrong, and the real one is a vacuous assertion rather than a missing one. Original text kept below the line for the record.
Actual mechanism
The spec DOES await the persist, at
tools/simulation/ui/specs/i18n.spec.ts:95-100:That assertion passed in the failing run — the failure came later, at line 156. So "the reload beat an un-awaited PUT" is refuted by the code.
What actually happens is that the listener matches any
PUT /me/language, and there is another one in flight:selector.selectOption("en")to normalise the starting language. That firesPUT {language:"en"}and is not awaited.waitForResponsefor any PUT to/me/language.selectOption(lang).LanguageSelector.tsx:44renders<select … disabled={busy}>, so Playwright's actionability check waits for theenflight to finish. During that wait theenresponse arrives — and satisfies the listener registered at step 3..ok()on the reset's response. It passes vacuously, while thelangPUT has only just been fired.localStorage.clear()+reload()at 134-135 can now outrun thelangPUT, so/mereturns no preference and the shell renders English — the observed failure.esis test #1 and pays cold start, which is what makes theenPUT slow enough to still be in flight at step 3. On a warm run it has already settled, the listener correctly matches thelangPUT, and everything holds — hence one failure in twenty runs.This is the #428 shape: a listener bound to the wrong request.
Fix
Bind the wait to the specific language, so the reset's response cannot satisfy it:
putMeLanguagesends{ language }(web/src/api/cluckwork.ts:420), so the payload is the discriminator. Ordering needs no extra handling: the select is inert during theenflight, so that PUT is always complete beforelangis selected.The vacuity is the real defect here — the assertion on line 100 exists precisely to prove the preference was persisted, and it can pass without ever observing that request.
Original report (mechanism since refuted)
Seen once on CI, on a docs-only PR whose diff cannot reach this job (#484 touches
AGENTS.mdand a comment indeploy/docker-compose.yml; the sim harness runs the self-containedtools/simulation/docker-compose.sim.yml). Same app bytes asmain, which passed an hour later — so the discriminator is timing, not code.Evidence
Failing job: run 31345832064 — 29 passed, 1 failed.
Passing job on
main, same code: run 31348597389.Durations for the same parameterized test:
es(test #1)tl(test #2)mainOn
mainthe whole test — switch, clear storage, reload, re-assert — finished in 2.3s, nowhere near the timeout. On the failing run the post-reload assertion burned the full 10s.The uploaded screenshot shows the post-reload page is healthy and entirely English: English nav, and the Language selector itself reading
English. Not a broken render, not a missing catalog key, not a crash — the app correctly showing English because no preference was persisted to find. The selector is enabled, so the flight had settled by then.Permissions are ruled out:
tlran the same body with the same Read-only persona in the same run and passed.Mechanism
web/src/session/LanguageSelector.tsx:33-38fires both calls and awaits neither:The local switch repaints immediately, so the test's first assertions (
Informesvisible,<html lang="es">) pass off device state alone. The test then does, attools/simulation/ui/specs/i18n.spec.ts:134-135:with nothing awaiting the in-flight
PUT /me/language. If the reload beats the PUT, nothing was persisted,/mereturns no language, and the shell renders English — the observed failure.esis test #1 and pays cold start (service worker install, cold container), which is why it is the one that loses.What is NOT proven
The PUT was never directly observed. The container-log dump is failure-only and tail-truncated (the failing run's dump covers only the last 19s, starting at 01:01:00, while the test ran ~01:00:21-33; the passing run dumped nothing at all), and the artifact upload carries only
video.webmand the screenshot —trace.zipis generated but not uploaded, and the trace is what holds the network log.So "the reload beat the PUT" vs "the PUT was issued and failed" is inferred from source plus timing, not observed. Same defect class and same fix either way, but the distinction is currently unfalsifiable on CI. See #NEXT for the artifact gap that made it so.
Proposed fix
LanguageSelector.tsx:44already renders<select … disabled={busy}>— the control is inert while the persist is in flight, deliberately (see the comment on line 32). So the test has a deterministic signal: wait for the select to re-enable before clearing storage and reloading. Does not weaken the assertion, and matches the established pattern for this shape (wait on the disabled-while-pending control rather than on a fixed delay).Related
Same family as #470 (arrange-phase probe that could not distinguish "not started" from "finished") and #428 (listener registered after the request it waits for): a test asserting state that depends on an un-awaited async write.