Skip to content

Don't let cooperative_wait's worker threads deadlock - #26

Merged
maleadt merged 1 commit into
mainfrom
tb/coopwait-robustness
Oct 1, 2026
Merged

maleadt merged 1 commit into
mainfrom
tb/coopwait-robustness

Conversation

@maleadt

@maleadt maleadt commented Oct 1, 2026

Copy link
Copy Markdown
Member

cooperative_wait hands blocking waits to a small pool of worker threads, so that the task waiting for the GPU doesn't block its thread. While hardening CUDA.jl's switch to it (JuliaGPU/CUDA.jl#3315), we found two ways in which those workers could still deadlock. This PR fixes both.

Workers no longer run finalizers

The worker threads are adopted by Julia, so they can run finalizers: when they trigger a collection, and when they release a lock with finalizers pending (e.g. when putting themselves back in the pool, right before waking up the waiting task). Finalizers of GPU objects can block for a long time, though. On CUDA, for example, these calls wait for all running kernels to finish (RTX 5080, with a 500 ms kernel running in the background):

cuMemFree                  454.8 ms
cuMemFreeHost              454.8 ms
cuMemHostUnregister        454.9 ms
cuMemAlloc                   0.1 ms
cuStreamDestroy              0.0 ms

So a worker that had just finished waiting for one stream could get stuck in a finalizer freeing memory, before telling the waiting task that its stream had completed. If the work running on the GPU depends on that task making progress (e.g. a kernel waiting for a flag the host sets), nothing ever completes.

Workers now permanently inhibit finalizers. Finalizers that become pending on a worker thread are run by regular Julia threads instead. The new test triggers a collection from within a worker. Without this change, the finalizer runs there:

Test Failed at test/synchronization.jl:154
  Expression: :foreign ∉ pools
   Evaluated: foreign ∉ [:foreign]

Objects that cannot be polled no longer wait for a worker

The pool has at most 4 workers, because drivers often spin while waiting, so every busy worker can occupy a CPU core. When all of them are busy, objects that can be polled (streams, events) are polled on the calling thread until a worker frees up. Objects that cannot be polled (e.g. an entire CUDA context) had to wait for one instead. That can deadlock: if the 4 busy workers are waiting for GPU work that only completes after the waiting task gets to continue, no worker ever becomes available.

Such waits now get an extra worker when all regular ones are busy. These overflow workers are kept in a separate pool that only serves objects that cannot be polled, so waits for pollable objects still use at most 4 workers. Without that separation, every overflow worker would also become available to pollable waits, and the number of workers busy waiting on pollable objects would grow with every burst. The new test keeps all regular workers busy with operations that never complete, and checks that a context-like wait still finishes. Before, it timed out:

Test Failed at test/synchronization.jl:130
  Expression: timedwait(() -> istaskdone(nonpolled), 30) === :ok
   Evaluated: timed_out === ok

Both tests pass with these changes on Julia 1.10, 1.12 and 1.13, as does CUDA.jl's test suite. This is a bug-fix release (3.3.1).

Workers no longer run finalizers. These may block, e.g., when freeing
GPU memory waits for the device to become idle, which would keep the
worker from notifying its waiter, and could deadlock if what the
finalizer waits for depends on that.

Objects that cannot be polled no longer wait for a worker when all of
them are busy, which could deadlock when the busy workers wait for
operations that only complete after this wait does. They use separate
overflow workers instead, so the number of workers that waits for
polled objects can occupy remains bounded.
@maleadt
maleadt merged commit 24dd303 into main Oct 1, 2026
10 checks passed
@maleadt
maleadt deleted the tb/coopwait-robustness branch October 1, 2026 08:19
maleadt added a commit to JuliaGPU/OpenCL.jl that referenced this pull request Oct 1, 2026
On GPUs, waits are handed to GPUToolbox's worker threads, which 3.3.1 keeps
from deadlocking: workers no longer run finalizers, which could block (e.g.,
freeing memory waits for the device) before the worker notifies its waiter
(JuliaGPU/GPUToolbox.jl#26).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant