Don't let cooperative_wait's worker threads deadlock - #26
Merged
Merged
Conversation
Workers no longer run finalizers. These may block, e.g., when freeing GPU memory waits for the device to become idle, which would keep the worker from notifying its waiter, and could deadlock if what the finalizer waits for depends on that. Objects that cannot be polled no longer wait for a worker when all of them are busy, which could deadlock when the busy workers wait for operations that only complete after this wait does. They use separate overflow workers instead, so the number of workers that waits for polled objects can occupy remains bounded.
maleadt
added a commit
to JuliaGPU/OpenCL.jl
that referenced
this pull request
Oct 1, 2026
On GPUs, waits are handed to GPUToolbox's worker threads, which 3.3.1 keeps from deadlocking: workers no longer run finalizers, which could block (e.g., freeing memory waits for the device) before the worker notifies its waiter (JuliaGPU/GPUToolbox.jl#26).
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
cooperative_waithands blocking waits to a small pool of worker threads, so that the task waiting for the GPU doesn't block its thread. While hardening CUDA.jl's switch to it (JuliaGPU/CUDA.jl#3315), we found two ways in which those workers could still deadlock. This PR fixes both.Workers no longer run finalizers
The worker threads are adopted by Julia, so they can run finalizers: when they trigger a collection, and when they release a lock with finalizers pending (e.g. when putting themselves back in the pool, right before waking up the waiting task). Finalizers of GPU objects can block for a long time, though. On CUDA, for example, these calls wait for all running kernels to finish (RTX 5080, with a 500 ms kernel running in the background):
So a worker that had just finished waiting for one stream could get stuck in a finalizer freeing memory, before telling the waiting task that its stream had completed. If the work running on the GPU depends on that task making progress (e.g. a kernel waiting for a flag the host sets), nothing ever completes.
Workers now permanently inhibit finalizers. Finalizers that become pending on a worker thread are run by regular Julia threads instead. The new test triggers a collection from within a worker. Without this change, the finalizer runs there:
Objects that cannot be polled no longer wait for a worker
The pool has at most 4 workers, because drivers often spin while waiting, so every busy worker can occupy a CPU core. When all of them are busy, objects that can be polled (streams, events) are polled on the calling thread until a worker frees up. Objects that cannot be polled (e.g. an entire CUDA context) had to wait for one instead. That can deadlock: if the 4 busy workers are waiting for GPU work that only completes after the waiting task gets to continue, no worker ever becomes available.
Such waits now get an extra worker when all regular ones are busy. These overflow workers are kept in a separate pool that only serves objects that cannot be polled, so waits for pollable objects still use at most 4 workers. Without that separation, every overflow worker would also become available to pollable waits, and the number of workers busy waiting on pollable objects would grow with every burst. The new test keeps all regular workers busy with operations that never complete, and checks that a context-like wait still finishes. Before, it timed out:
Both tests pass with these changes on Julia 1.10, 1.12 and 1.13, as does CUDA.jl's test suite. This is a bug-fix release (3.3.1).