Skip to content

fix(e2e): session-races 'explicit login beats pending bootstrap refresh' races goto() vs waitForRequest — fails on CI's bundled Chromium #428

Description

@mforce

What's failing

tools/simulation/ui/specs/session-races.spec.ts:232 — #310 session races › an explicit login beats a pending bootstrap refresh, and the newer login wins.

Reported failing run: https://github.com/mforce/cluckwork/actions/runs/31077917484/job/92539914682

Since when

Bisected via E2E-smoke run history (this workflow is workflow_dispatch-only, so nothing runs it automatically — see #370):

  • Last green: 2026-08-02 21:47 @ 72e84b59d
  • First red: 2026-08-03 09:15 @ ef9a64ba7 — same test, same error, every run since (today's run on feat/273-... included).
  • Of the 5 commits in that window, only b39e8fb ("Credential epoch: per-request revocation checks", Credential epoch: per-request revocation checks #399) touches web/src/api/client.ts or web/src/auth/AuthContext.tsx.

However — see Root cause below — #399 is not where the defect lives. It most likely just shifted request timing (extra per-request work) enough to flip a pre-existing race from "usually passes" to "reliably fails" in CI's environment specifically.

Root cause — reproduced and nailed down locally

Not an application bug. It's a Playwright test race in the spec itself, and it's browser-build-dependent, which is why it was invisible locally:

  • npx playwright test with system Chromium → passes, every time (5.8s).
  • CLUCKWORK_E2E_BUNDLED_BROWSER=1 (Playwright's own downloaded Chromium, 151.0.7922.34 — what CI uses) → fails deterministically, every time, with the exact CI error.

Trace of the failing run (--trace on), full chronological network log, decisive part:

900.08  GET  200              /login                          <- test's page.goto("/login")
903-951 ...  200              (css/js/font assets loading)
952.63  POST -1 ERR_ABORTED   /api/v1/auth/refresh             <- the bootstrap refresh — ABORTED
```//
followed by the test hanging 45s on `page.waitForRequest(...)` because that's the *only* refresh attempt this page load makes — `restoreSession()` doesn't retry on failure — and the listener registered after `page.goto()` resolves never sees a request that already happened.

**The bug:** `session-races.spec.ts:256-257`

```ts
await page.goto("/login");
await page.waitForRequest((r) => r.url().includes("/api/v1/auth/refresh"));

page.goto() with the default waitUntil: "load" doesn't resolve until the browser's load event — which waits on every subresource, including the 7 web-font requests this page issues. On a slow-enough asset load (CI's bundled Chromium; not local system Chromium), React mounts and fires AuthContext's bootstrap restoreSession() fetch before load fires — i.e., before the test's await page.goto(...) line even returns control. By the time the test reaches the waitForRequest(...) line, the request has already happened (and in this exact browser build, gets cancelled — net::ERR_ABORTED — evidently as a side effect of still being CDP-intercepted via the test's own page.route(...) when the navigation's load state settles). Either way, the listener is registered too late to observe it.

This is the exact anti-pattern Playwright's own docs warn about for goto + a same-navigation wait: register the waiter before or concurrently with the navigation, not after.

The sibling test two tests up (session-races.spec.ts:142, passes reliably) has the same superficial "await X, then await waitForRequest" shape but is safe, because it provokes the refresh via nav.link(...).click() — a client-side route change, not a full page.goto() — so it's never racing a load event against 7 font requests. Grepped the rest of specs/ — this page.goto() + same-navigation waitForRequest() combination is otherwise unique to this one test; not a systemic pattern in the suite.

Confirmed fix

Standard Playwright pattern — register the listener and the navigation together:

const [request] = await Promise.all([
  page.waitForRequest((r) => r.url().includes("/api/v1/auth/refresh")),
  page.goto("/login"),
]);

Verified locally against CLUCKWORK_E2E_BUNDLED_BROWSER=1 (CI's exact browser) — 4/4 green, ~5.8–6.2s each, vs. deterministic 45s-timeout failure before. (A narrower page.goto("/login", { waitUntil: "commit" }) change also fixes it in isolation, consistent with the same root cause, but Promise.all is the correct fix — it doesn't depend on load-timing margins at all.)

Scope

One test file, tools/simulation/ui/specs/session-races.spec.ts, lines ~256-257. No application code change needed.

Activity

  1. KNullerr commented on Aug 6, 2026

    @KNullerr

    Hi Mforce, I saw your issue opened today. The problem was caused by the waitForRequest listener being registered after page.goto(), leading to a race condition when loading resources, especially in slower CI environments. The solution is to register the listener simultaneously with navigation to ensure no event leaks. I've written the exact patch for you:
    Modify the code in session-races.spec.ts according to this structure:

    const [request] = await Promise.all([
    page.waitForRequest((r) => r.url().includes("/api/v1/auth/refresh")),
    page.goto("/login"),
    ]);

    . If you'd like to test it and need help securing the entire system to avoid further crashes, I can fix it in an hour. I only ask for an appropriate compensation, preferably payable in crypto only once the issue is resolved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions