Skip to content

Fix intermittent crashes at thread and interpreter teardown - #4110

Closed
dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:teardown
Closed

dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:teardown

Conversation

@dhiltgen

Copy link
Copy Markdown
Contributor

Proposed changes

MLX sometimes dies at thread or process teardown with "terminate called without an active exception" or SIGSEGV, on Linux/Windows. The general pattern: a destructor reachable from __call_tls_dtors() must not acquire the GIL, nor throw, nor drop a Python reference.

Acquiring the GIL is fatal because take_gil() answers a request from a non-finalizing thread, while another thread finalizes, by calling PyThread_exit_thread(); the forced unwind then crosses the noexcept destructor frame.

Throwing is equally fatal: CUDA destructors ran CHECK_*_ERROR on teardown calls that legitimately fail -- every Windows run with two threads died.

Adds python/tests/test_teardown.py. The failure kills the process, so the cases run in subprocesses and repeat; both fail without this change.

Checklist

Put an x in the boxes that apply.

  • I have read the CONTRIBUTING document
  • I have run pre-commit run --all-files to format my code / installed pre-commit prior to committing changes
  • I have added tests that prove my fix is effective or that my feature works
  • I have updated the necessary documentation (if needed)

MLX sometimes dies at thread or process teardown with "terminate called
without an active exception" or SIGSEGV, on Linux/Windows.  The general
pattern: a destructor reachable from __call_tls_dtors() must not acquire
the GIL, nor throw, nor drop a Python reference.

Acquiring the GIL is fatal because take_gil() answers a request from a
non-finalizing thread, while another thread finalizes, by calling
PyThread_exit_thread(); the forced unwind then crosses the noexcept
destructor frame.

Throwing is equally fatal: CUDA destructors ran CHECK_*_ERROR on teardown
calls that legitimately fail -- every Windows run with two threads died.

Adds python/tests/test_teardown.py. The failure kills the process, so the
cases run in subprocesses and repeat; both fail without this change.

@zcbenz zcbenz left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Our current recommended pattern is to let users call clear_streams at the end of threads and the process, so they can do cleanup before cuda and python shutdown, which is something that we could not do on the library side. Have you tried that? We should have a multi-threading doc for all these APIs though.

TEST_CASE("test new stream in threads") {
std::vector<std::thread> threads;
for (int i = 0; i < 1; ++i) {
threads.emplace_back([]() {
auto s = new_stream(default_device());
eval(arange(10, s));
clear_streams();
});
}
for (auto& t : threads) {
t.join();
}
}

@dhiltgen

Copy link
Copy Markdown
Contributor Author

I can close this PR, these are just crashes that I was noticing while working on other improvements.

@zcbenz zcbenz closed this Aug 11, 2026
@dhiltgen
dhiltgen deleted the teardown branch August 11, 2026 23:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants