Use GPUToolbox's cooperative_wait for synchronization - #3315
Merged
Merged
Conversation
The worker threads used for nonblocking synchronization were assigned to tasks round-robin and then kept, so a long synchronization delayed unrelated tasks sharing its worker. GPUToolbox now provides this functionality as a shared, reworked implementation: a pool of workers that each handle one request at a time, with interrupts deferred until the operation has completed. Device synchronization also no longer polls the legacy stream first, which does not cover non-blocking streams, and then blocked the thread in cuCtxSynchronize.
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
`used_memory` and `cached_memory` only need the device, but fetched it through `active_state()`, which creates the task's default stream. They are called by `maybe_collect` right before synchronizing, and creating a stream can block the thread until the GPU is idle when the driver needs to grow its pool of streams.
The default and legacy streams have no context, so handing them to a worker thread failed. Pass the caller's context along instead. The per-thread stream is specific to the calling thread, so wait for an event recorded on it instead. Also don't synchronize again on the calling thread when polling found the object to have completed: that could block the thread on work that another task submitted in the meantime. `isdone` already reports errors, and a successful query counts as synchronization.
The tests relied on kernels running for a fixed amount of time, which failed on CI when the thread was blocked creating a stream: the driver waits for running kernels to finish when it grows its pool of streams. Instead, keep the GPU busy until another task opens a gate, with a timeout the tests check for, and set up all streams and kernels before. Also cover the default, legacy and per-thread streams.
GPUToolbox 3.3.1 fixes two ways in which cooperative_wait's worker threads could deadlock: a worker running a finalizer that waits for the GPU (e.g. freeing memory) before notifying its waiter, and a wait for an object that cannot be polled, like the context in device_synchronize, waiting for a worker while all of them were busy.
maleadt
added this pull request to stack #3319
October 1, 2026 11:24
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #3315 +/- ##
==========================================
+ Coverage 86.02% 86.10% +0.08%
==========================================
Files 187 187
Lines 19042 18971 -71
==========================================
- Hits 16380 16335 -45
+ Misses 2662 2636 -26 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
maleadt
added a commit
that referenced
this pull request
Oct 3, 2026
* Fix waiting for GPU work before releasing wrapped memory. The release of wrapped system memory still called `nonblocking_synchronize`, which #3315 removed. Waiting failed with an UndefVarError that was only logged, after which the memory was unregistered and the Array released while the GPU could still be using it. Use `cooperative_wait` instead. Also wait and unregister with a relaxed capture mode: freeing a wrapper while another stream is being captured, e.g. by the task doing the freeing, otherwise invalidated that capture. * Don't wait for a whole context while capturing. Releasing wrapped memory last used on a special or destroyed stream waits for the entire context, which is prohibited while any of its streams is being captured, regardless of the capture mode. Postpone that until captures made with `capture` have ended.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CUDA.jl's nonblocking synchronization, which waits for the GPU on a worker thread so the calling task doesn't block its thread, now lives in GPUToolbox as
cooperative_wait(JuliaGPU/GPUToolbox.jl#23), shared with oneAPI.jl, OpenCL.jl and KernelAbstractions. This PR switches CUDA.jl over to it, which removes ~200 lines here and fixes a few problems with the old implementation along the way:device_synchronize()could block the thread. It first polled the legacy stream, which doesn't cover streams created withSTREAM_NON_BLOCKING, and if that said "done", it calledcuCtxSynchronizeon the calling thread, blocking it until those other streams finished. It now always waits on a worker.default_stream()andlegacy_stream()have no context, so the worker failed. They now use the caller's context.per_thread_stream()is only meaningful on the calling thread, so it's synchronized through an event recorded on it.The polling heuristic is unchanged, so latency is the same. Comparing both implementations in one process, alternating between them, the median overhead of a launch plus
synchronize()minus the kernel duration on an RTX 5080:(These were measured on a heavily loaded machine, which adds ~2 µs across the board. For longer kernels, both versions are dominated by waking up the waiting thread and were within noise of each other.)
The new tests don't depend on timing: the GPU is kept busy by a kernel that waits until another task on the same thread opens a gate, and the tests check that this happened before the kernel's timeout. They cover other tasks running during
synchronize(stream),synchronize(event),device_synchronize()and synchronization of the default streams, and a long synchronization not delaying shorter ones from other tasks. Thedevice_synchronize()test and the last one fail on master, the default stream ones failed with the first version of this PR.This requires GPUToolbox 3.3.1, which fixes two ways in which the worker threads could deadlock that came up while working on this PR (JuliaGPU/GPUToolbox.jl#26): a worker could get stuck running a finalizer that waits for the GPU (e.g. freeing memory) before waking up its waiting task, and
device_synchronize(), which cannot poll, had to wait for a worker when all of them were busy.