Skip to content

test:unit:main hangs/flakes on macOS: applicationSpawn.test.js never finishes, EntryHandler.test.js first test times out ~4/6 #2538

Description

@Devin-Holland

Summary

Two test:unit:main files cannot be relied on locally on macOS, measured on pristine origin/main (1a5067e1b) with Node 24.19 and a freshly installed node_modules (including @harperfast/rocksdb-js-darwin-arm64@2.8.0):

  • unitTests/components/applicationSpawn.test.js hangs indefinitely after the nonInteractiveSpawn process-group tests. The last test to report is "accepts a Windows taskkill miss only when the process tree is independently gone"; mocharc's timeout: 0 means the run never ends, so npm run test:unit:main never prints a summary (it was killed at 1700 s twice; the file alone never finished inside 240 s).
  • unitTests/components/EntryHandler.test.js — its first test, "should instantiate and emit events for adding and removing files and directories", times out in ~4 of 6 runs (--timeout 20000 to make the hang visible; 39 of 40 tests pass). The other 39 pass every time, so this looks like the chokidar/FSEvents warm-up race on a fresh process rather than a logic failure.

Both files are identical between main and the branch I was verifying (#2377), and the failure rates matched exactly on both (4/6 and 4/6), so this is the platform, not a regression. Linux CI is green on the same files.

Impact

Anyone running the main gate on a Mac gets no summary at all (the hang) and cannot tell a real failure from the flake. Local workaround: exclude the two files (--exclude unitTests/components/applicationSpawn.test.js --exclude unitTests/components/EntryHandler.test.js) and rely on CI for them.

Suggested fix shape

  • applicationSpawn.test.js: bound the process-tree waits on non-Linux (the Linux /proc scan paths are platform-specific by design; on Darwin the wait appears to have no exit condition), or skip the Linux-only group with this.skip() on process.platform !== 'linux'.
  • EntryHandler.test.js: wait for the chokidar ready event (or a first no-op event) before touching the tree in the first test, and give the file a real mocha timeout so a hang fails instead of stalling the gate.

Activity

  1. added theissue type on Sep 9, 2026
  2. kriszyp commented on Sep 14, 2026

    @kriszyp
    Member

    Counter-example to "Linux CI is green on the same files" — EntryHandler.test.js hung on Linux CI, Node v22, not just macOS.

    harper#2498 → run 34803736178, job 103851368764. The last line the suite printed was the EntryHandler describe header, and nothing after it:

    2026-09-14T03:47:44.2499114Z   EntryHandler
    2026-09-14T04:02:22.7507795Z ##[error]The action 'Run tests' has timed out after 15 minutes.
    

    14m38s between the header and the step timeout, with no test result in between — the same "first test never returns" shape this issue describes at ~4/6 on Darwin, reproduced once on ubuntu-latest. Node v22; v24, v26 and Windows v24 passed the identical head.

    Two things that follow:

    • The scope line can probably drop "macOS" — if it is the chokidar warm-up race, Linux is rarer but not immune, and a rare hang in CI is worse than a frequent one locally because it consumes the whole step.
    • Because test:unit:all chains four suites into one 15-minute step, this hang did not cost the EntryHandler result — it cost every later suite's result too (apitests, resources, lmdb never ran). That blast radius is filed separately as the step-headroom issue, which links back here.

    Found while maintaining #2498; not a regression from that branch (the file is untouched by it, and main's own v22 run was green the same night).

    — Claude Opus 5

  3. cb1kenobi commented on Sep 29, 2026

    @cb1kenobi
    Member

    Reproduced independently today on a clean origin/main worktree at 80f6d1c38 (verified as a real ancestor commit) — same file, same shape, not branch-specific.

    Mechanism, traced in the current applicationSpawn.test.js: the last test, terminates a detached process tree when its owning worker is force-terminated, spawns unitTests/components/fixtures/processGroupHarness.js as a real child process and does await once(harness, 'close') with no escape hatch. The harness starts a worker (manageThreads.startWorker), which spawns a further descendant via nonInteractiveSpawn; on the worker's ready message it calls worker.terminate() and prints terminated. If that termination doesn't fully release the process/thread, close never fires — the awaited promise, and mocha's entire run, hang.

    .mocharc.json's "timeout": 0 (confirmed on origin/main) is what turns this into a silent wedge instead of a failure: no test ever times out, so there's no error and no log — just an indefinite stall with no summary line.

    New: orphan accumulation. Three parentless processGroupHarness.js processes were found alive at once on one machine today, the oldest 11h54m old, left behind by earlier interrupted runs — the hang doesn't end when the parent session does, it just detaches. They were killed; until someone notices and does the same, they degrade every later run on that machine.

    Cost: this isn't just "the suite is slow." A gate that cannot complete locally is a gate that gets reported as unrun, or worse — it already produced one overclaimed test result that had to be corrected on a live PR today.

    Not independently reproduced: a third report of the same suite stalling in the OCSP/CRL certificate tests once applicationSpawn.test.js is excluded. Flagging as unconfirmed, not restating it as fact.

    — Claude Sonnet 5

  4. added this to the v5.3 milestone on Sep 29, 2026
  5. modified the milestones: v5.3, v5.4 on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P1

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions