Skip to content

E2E: manager write-off spec fails on a repeatedly-reused simulation fixture #506

Description

@mforce

What happened

During #504's mutation runs, mutation-check.sh aborted on a RED baseline:

1 failed
  [chromium] › specs/manager.spec.ts:183:3 › Manager › writes off lost stock
    from its own lot without touching the entry
1 skipped
32 passed

The spec passes in isolation immediately afterwards, and the full suite is green again after reset.sh. Nothing in #504 touches that spec or any code it exercises.

Evidence, and its limits

One occurrence, after roughly ten consecutive full-suite runs against a single seeded fixture (baseline + restore for each of several mutation runs). The Playwright artifacts were cleaned by the next passing run, so I have the failure line and nothing more — no error context, no screenshot. I could not reproduce it deliberately, and I did not try hard, because a reset made it go away and the work it was blocking was unrelated.

So this is a report, not a diagnosis.

Hypothesis

specs/manager.spec.ts:183 creates its own flock and lot per run, sized

const eggs = 41 + (Date.now() % 58);   // "41–98, effectively unique per run"

and then finds the lot with a row filter on today's date and that exact count:

.filter({ has: page.getByRole("cell", { name: today, exact: true }) })
.filter({ has: page.getByRole("cell", { name: String(eggs), exact: true }) })
.first()

Two things about that survive a single run but not many:

  1. eggs is 1-in-58, not unique. Ten runs on one fixture day make a collision likely. .first() then silently selects an earlier run's lot, which may already be written off.
  2. Lots accumulate. Every run adds a same-day lot to the same grade. If the lot list paginates, the target can eventually fall outside the rows the filter can see.

Both are invisible on a fresh fixture, which is the normal case, and both get steadily more likely the more the suite is re-run — exactly the pattern of a local mutation session.

Why it is worth an issue rather than a shrug

The repo's rule is that a flake is a finding. This one has a specific edge: it aborts the mutation harness, and it does so at the baseline, which is the one place the harness is designed to stop rather than continue. Anyone running npm run mutation more than a few times without a reset will hit it and lose the run, and the failure names a spec unrelated to whatever they were checking.

Possible fixes, in order of preference

  • Make the lot genuinely unique per run — the flock name already uses Date.now() in full; the egg count could key off the same value rather than % 58. The comment explains the 41–98 range is chosen to stay disjoint from the flow spec's 38–40, so the constraint is a range, not a modulus.
  • Or scope the row filter by the flock this test just created, rather than by date + count.
  • Or have mutation-check.sh say so when the baseline fails on a spec unrelated to the mutants being run, and suggest reset.sh.

Not doing it in #504

That PR is about live regions under a modal; this is a pre-existing fragility in a different spec, surfaced by it rather than caused by it. Folding a fix in would widen the diff and mix two unrelated arguments.

Activity

  1. mforce commented on Aug 11, 2026

    @mforce
    OwnerAuthor

    Fixed in 24fc61b5, on PR #504 — the PR that surfaced it, rather than as a follow-up.

    Correcting this issue's own diagnosis first. The text above blames the 1-in-58 count collision. That is at most half of it, and I wrote it from a symptom rather than from evidence. Two further hypotheses I raised or implied — the 50-row paging ceiling and the #464 prefill wipe — were also wrong, and I checked both: today's lots all sit on page 1 (28 of them, with paging adding only older dates), and the prefill synchronisation is working.

    What actually breaks it. Two filter({ has: cell }) clauses naming the same text are satisfied by the same cell. So while produced equals available, the row filter degenerates into "a today row containing that number anywhere" — and it matched a lot whose AVAILABLE happened to equal our produced count. The spec then wrote off a stranger's lot, taking it from -2 to -4, and left its own untouched for the following assertion to miss. Rows with a gap of 4 in a well-used fixture are that mistake's fingerprints.

    Collision is the enabling condition — on a fresh fixture nothing else on screen carries that number — which is why it only bites after several runs and why a reset makes it disappear.

    The fix. The row is addressed by the pair (produced, available) positionally — produced is td:nth-child(2), available is td:nth-child(3) — and re-addressed at its new balance after the write-off rather than reusing one locator whose predicate the write-off invalidates.

    Evidence. 12 consecutive runs against a deliberately accumulated fixture pass, where the previous version failed 3 of 12. Full suite green.

    What I would take from it beyond the fix: the DOM dumped at the moment of failure settled this in one run, after three plausible theories had each cost more than that. That should have been the first move, not the fourth.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions