Repository navigation
Terminate workers blocked in Atomics.wait() when recording - #175
Merged
Merged
Conversation
A worker blocked in Atomics.wait() and terminated while recording never stopped, so the process never exited. The stop watchdog (#173) invalidated the recording after 5s and terminated the worker's execution, which woke the futex wait, but FutexEmulation::WaitSync handles interrupts with events disallowed (#170), and StackGuard::HandleInterrupts ignores interrupts while events are disallowed (#173, ported from replayio/chromium-v8#30). The termination stayed pending and the waiter went back to waiting on the futex emulation's condition variable, with the parent idle in its event loop, waiting for the worker's exit. WaitSync now handles a termination that HandleInterrupts left pending. Only another thread can terminate a thread blocked there, which already invalidates the recording (StackGuard::RequestInterrupt), so this doesn't change what a usable recording or its replay does. Other interrupts are still left for the next stack check, and without recording HandleInterrupts has already handled the termination, so the check is a no-op. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When recording or replaying, Worker::Exit called from another thread posts the stop to the worker's event loop and starts the stop watchdog (#173), so that the worker stops at a point the replay can reproduce. When the parent exits, through process.exit() or the end of its event loop, the recording is already finished (RecordReplayFinishRecording runs before stop_sub_worker_contexts), so there is nothing left to reproduce: a worker that never returned to its event loop, e.g. one blocked in Atomics.wait(), held up the exit for the watchdog's 5s for no benefit. Worker::Exit now stops the worker right away once the recording is finished, as without recording. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Andarist
force-pushed
the
andarist/worker-terminate-hang
branch
from
October 7, 2026 12:27
ba44be5 to
b9189dc
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A worker blocked in
Atomics.wait()and terminated while recording never stopped, so the process never exited. The stop watchdog (#173) invalidated the recording after 5s and terminated the worker's execution, which woke the futex wait. ButFutexEmulation::WaitSynchandles interrupts with events disallowed (#170), andStackGuard::HandleInterruptsignores interrupts while events are disallowed (#173, ported from replayio/chromium-v8#30). The termination stayed pending and the waiter went back to waiting on the futex emulation's condition variable, with the parent idle in its event loop, waiting for the worker's exit.WaitSyncnow handles a termination thatHandleInterruptsleft pending. Only another thread can terminate a thread blocked there, which already invalidates the recording (StackGuard::RequestInterrupt), so this doesn't change what a usable recording or its replay does. Other interrupts are still left for the next stack check, and without recordingHandleInterruptshas already handled the termination, so the check is a no-op.Once the recording is finished,
Worker::Exitnow stops workers right away instead of posting the stop to their event loop and starting the stop watchdog. When the parent exits, throughprocess.exit()or the end of its event loop, the recording is already finished beforestop_sub_worker_contextsruns, so a deterministic stop gains nothing: a worker that never returned to its event loop, e.g. one blocked inAtomics.wait(), held up the exit for the watchdog's 5s. With a stuck worker, exiting now takes about as long as on stock node 16.Tested with https://github.com/replayio/backend/pull/13450 and its
test/node-recording/tests/worker-threads/terminate-stuck/