Skip to content

test(wal): wait for async writes instead of racing the writer goroutine - #566

Merged
xe-nvdk merged 1 commit into
mainfrom
fix/wal-purge-inactive-test-flake
Jul 29, 2026
Merged

xe-nvdk merged 1 commit into
mainfrom
fix/wal-purge-inactive-test-flake

Conversation

@xe-nvdk

@xe-nvdk xe-nvdk commented Jul 29, 2026

Copy link
Copy Markdown
Member

Follow-up to #565, where I flagged TestPurgeInactive_ConcurrentWithWrite failing under Linux containers.

It is a test bug, not a product bug — and not flaky

I measured it rather than assuming:

Platform Result
Linux (container), 20 runs 19 failures
macOS, 20 runs 0 failures

Nearly deterministic on Linux. "Flaky" was the wrong label — the test relies on a race the writer goroutine has no particular reason to win, and macOS's scheduler simply happens to let it.

Root cause

Append is asynchronous. It enqueues onto entryChan and returns (wal.go:444-453); a background writer goroutine performs the file write and only then increments TotalEntries (wal.go:300).

The test closed its producer channel once the 100 Append calls had been enqueued, then immediately asserted TotalEntries != 0. <-done says nothing about whether the writer ran.

I instrumented the actual values to confirm rather than infer:

TotalEntries immediately after <-done: 0
DroppedEntries:                        0
TotalEntries after 300ms settle:       100

Nothing was dropped and every entry was written. The WAL behaved correctly throughout; the test just looked too early.

Fix

Added waitForEntries, which polls until the writer has durably recorded the entries, or fails after 5s with a diagnostic (TotalEntries=… dropped=…) so a future failure says why.

Also replaced two fixed time.Sleep(50 * time.Millisecond) calls in TestPurgeInactive_ActiveFileNotDeleted. Those have the identical fragility with a larger margin — a fixed sleep only sets how often the race is lost, and on a loaded CI box or a slow container it can be lost too.

Verification

before after
TestPurgeInactive_ConcurrentWithWrite, Linux ×20 19 FAIL 0 FAIL
Full WAL suite, Linux container FAIL PASS (first time)
Full WAL suite, darwin PASS PASS
darwin -race ×10 PASS PASS

Test-only change — one file, no product code touched, so no release-notes entry.

Note the cross-compiled -race binary can't be built here (needs a Linux C toolchain), so the container runs are non-race builds. Darwin -race covers the detector side.

TestPurgeInactive_ConcurrentWithWrite failed 19 out of 20 runs on Linux
while passing 20 out of 20 on macOS. It is a test bug, not a product bug.

Append is asynchronous: it enqueues onto entryChan and returns, while a
background goroutine performs the file write and only then increments
TotalEntries. The test closed its producer channel once the 100 Appends
had been ENQUEUED and immediately asserted TotalEntries != 0, which races
the writer goroutine. macOS's scheduler happened to let the writer run;
Linux did not.

Instrumented the actual values to confirm rather than infer:

  TotalEntries immediately after <-done: 0
  DroppedEntries:                        0
  TotalEntries after 300ms settle:       100

Nothing was dropped and every entry was written — the test simply looked
too early.

Replaced the immediate read with waitForEntries, which polls until the
writer has recorded the entries or fails with a diagnostic after 5s. The
same fix replaces two fixed time.Sleep(50ms) calls in
TestPurgeInactive_ActiveFileNotDeleted, which have the identical
fragility with a larger margin.

Verified in a Linux container: the target test goes from 19/20 failures
to 0/20, and the full WAL suite now passes there for the first time.
Darwin stays green, including under -race.
@xe-nvdk
xe-nvdk merged commit 4a4e3ba into main Jul 29, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant