Skip to content

Use GPUToolbox's cooperative_wait for synchronization - #3315

Merged
maleadt merged 5 commits into
mainfrom
tb/coopsync
Oct 2, 2026
Merged

maleadt merged 5 commits into
mainfrom
tb/coopsync

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

CUDA.jl's nonblocking synchronization, which waits for the GPU on a worker thread so the calling task doesn't block its thread, now lives in GPUToolbox as cooperative_wait (JuliaGPU/GPUToolbox.jl#23), shared with oneAPI.jl, OpenCL.jl and KernelAbstractions. This PR switches CUDA.jl over to it, which removes ~200 lines here and fixes a few problems with the old implementation along the way:

  • A long synchronization could hold up unrelated tasks. Each task was assigned one of 4 worker threads, round-robin, and then kept it. Tasks sharing a worker therefore waited on each other: with 9 tasks synchronizing their own streams, one running 20 ms kernels, the tasks that shared its worker saw a p90 overhead of 18.8 ms on 1 ms kernels, against ~25 µs for the others. GPUToolbox's workers serve one request at a time and are handed out to whichever task needs one.
  • device_synchronize() could block the thread. It first polled the legacy stream, which doesn't cover streams created with STREAM_NON_BLOCKING, and if that said "done", it called cuCtxSynchronize on the calling thread, blocking it until those other streams finished. It now always waits on a worker.
  • Other synchronizations could block the thread too. When polling found the stream or event to have completed, it was synchronized again on the calling thread, which blocks if another task submitted work to it in the meantime. That second synchronization isn't needed: a successful query reports errors and counts as synchronization. And the GC heuristic that runs before waiting created the task's stream, which blocks until the GPU is idle when the driver needs to grow its pool of streams.
  • The default streams could not be synchronized on a worker. default_stream() and legacy_stream() have no context, so the worker failed. They now use the caller's context. per_thread_stream() is only meaningful on the calling thread, so it's synchronized through an event recorded on it.
  • Interrupting a synchronization could return too early, e.g. while a copy from host memory was still in flight. A Ctrl-C now takes effect once the operation has completed.

The polling heuristic is unchanged, so latency is the same. Comparing both implementations in one process, alternating between them, the median overhead of a launch plus synchronize() minus the kernel duration on an RTX 5080:

kernel blocking before after
0 µs 6.6 µs 6.8 µs 6.8 µs
50 µs 6.7 µs 7.0 µs 7.0 µs
100 µs 7.3 µs 7.7 µs 7.8 µs

(These were measured on a heavily loaded machine, which adds ~2 µs across the board. For longer kernels, both versions are dominated by waking up the waiting thread and were within noise of each other.)

The new tests don't depend on timing: the GPU is kept busy by a kernel that waits until another task on the same thread opens a gate, and the tests check that this happened before the kernel's timeout. They cover other tasks running during synchronize(stream), synchronize(event), device_synchronize() and synchronization of the default streams, and a long synchronization not delaying shorter ones from other tasks. The device_synchronize() test and the last one fail on master, the default stream ones failed with the first version of this PR.

This requires GPUToolbox 3.3.1, which fixes two ways in which the worker threads could deadlock that came up while working on this PR (JuliaGPU/GPUToolbox.jl#26): a worker could get stuck running a finalizer that waits for the GPU (e.g. freeing memory) before waking up its waiting task, and device_synchronize(), which cannot poll, had to wait for a worker when all of them were busy.

The worker threads used for nonblocking synchronization were assigned
to tasks round-robin and then kept, so a long synchronization delayed
unrelated tasks sharing its worker. GPUToolbox now provides this
functionality as a shared, reworked implementation: a pool of workers
that each handle one request at a time, with interrupts deferred until
the operation has completed.

Device synchronization also no longer polls the legacy stream first,
which does not cover non-blocking streams, and then blocked the thread
in cuCtxSynchronize.
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: c05ddfa Previous: af30d2d Ratio
array/accumulate/Float32/1d 99113 ns 99332 ns 1.00
array/accumulate/Float32/dims=1 72782 ns 72740 ns 1.00
array/accumulate/Float32/dims=1L 1588276 ns 1588667 ns 1.00
array/accumulate/Float32/dims=2 138133 ns 138016 ns 1.00
array/accumulate/Float32/dims=2L 654797 ns 654753 ns 1.00
array/accumulate/Int64/1d 117701 ns 118333 ns 0.99
array/accumulate/Int64/dims=1 77237 ns 76632 ns 1.01
array/accumulate/Int64/dims=1L 1699687 ns 1698176 ns 1.00
array/accumulate/Int64/dims=2 151027 ns 150013 ns 1.01
array/accumulate/Int64/dims=2L 987924 ns 986440 ns 1.00
array/broadcast 15648 ns 15515 ns 1.01
array/broadcast launch 6930.5 ns 6817.6 ns 1.02
array/construct 884.96 ns 877.0204081632653 ns 1.01
array/copy 16252 ns 16177 ns 1.00
array/copyto!/cpu_to_gpu 209116 ns 208346 ns 1.00
array/copyto!/gpu_to_cpu 241036 ns 240439 ns 1.00
array/copyto!/gpu_to_gpu 9148 ns 10126.333333333334 ns 0.90
array/iteration/findall/bool 130030 ns 131839 ns 0.99
array/iteration/findall/int 139283 ns 143795 ns 0.97
array/iteration/findfirst/bool 67303 ns 69501 ns 0.97
array/iteration/findfirst/int 68639 ns 71337 ns 0.96
array/iteration/findmin/1d 58941 ns 61367 ns 0.96
array/iteration/findmin/2d 96089 ns 96638 ns 0.99
array/iteration/logical 180887 ns 180714 ns 1.00
array/iteration/scalar 57948 ns 61462 ns 0.94
array/permutedims/2d 44814 ns 45696 ns 0.98
array/permutedims/3d 47140 ns 47162 ns 1.00
array/permutedims/4d 48655 ns 47964 ns 1.01
array/random/rand/Float32 11349 ns 11296 ns 1.00
array/random/rand/Int64 18351 ns 18630 ns 0.99
array/random/rand!/Float32 7777.75 ns 7761.5 ns 1.00
array/random/rand!/Int64 16401 ns 16176 ns 1.01
array/random/randn/Float32 32201 ns 32364 ns 0.99
array/random/randn!/Float32 23942 ns 23217 ns 1.03
array/reductions/mapreduce/Float32/1d 32082 ns 32174 ns 1.00
array/reductions/mapreduce/Float32/dims=1 37359 ns 37677 ns 0.99
array/reductions/mapreduce/Float32/dims=1L 50928 ns 51036 ns 1.00
array/reductions/mapreduce/Float32/dims=2 55120 ns 55166 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 67142 ns 67178 ns 1.00
array/reductions/mapreduce/Int64/1d 38560 ns 39967 ns 0.96
array/reductions/mapreduce/Int64/dims=1 40438 ns 40335 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 88289 ns 88397 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57691 ns 57744 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 83627 ns 83232 ns 1.00
array/reductions/reduce/Float32/1d 32076 ns 32541 ns 0.99
array/reductions/reduce/Float32/dims=1 37510 ns 37412 ns 1.00
array/reductions/reduce/Float32/dims=1L 50642 ns 50636 ns 1.00
array/reductions/reduce/Float32/dims=2 55357 ns 55121 ns 1.00
array/reductions/reduce/Float32/dims=2L 67581 ns 67277 ns 1.00
array/reductions/reduce/Int64/1d 38891 ns 40164 ns 0.97
array/reductions/reduce/Int64/dims=1 40453 ns 40300 ns 1.00
array/reductions/reduce/Int64/dims=1L 88644 ns 88417 ns 1.00
array/reductions/reduce/Int64/dims=2 57741 ns 57523 ns 1.00
array/reductions/reduce/Int64/dims=2L 83688 ns 83260 ns 1.01
array/reverse/1d 16739 ns 16957 ns 0.99
array/reverse/1dL 69567 ns 69738 ns 1.00
array/reverse/1dL_inplace 67232 ns 67182 ns 1.00
array/reverse/1d_inplace 8436.666666666666 ns 8408 ns 1.00
array/reverse/2d 20022 ns 19996 ns 1.00
array/reverse/2dL 73325 ns 73368 ns 1.00
array/reverse/2dL_inplace 67108 ns 67238 ns 1.00
array/reverse/2d_inplace 9595 ns 9746 ns 0.98
array/sorting/1d 2655739 ns 2640430 ns 1.01
array/sorting/2d 1018129 ns 1018269 ns 1.00
array/sorting/by 3171802 ns 3159862 ns 1.00
cuda/synchronization/context/auto 6953.5 ns 994.8 ns 6.99
cuda/synchronization/context/blocking 796.8461538461538 ns 779.196261682243 ns 1.02
cuda/synchronization/context/nonblocking 6893.75 ns 5818 ns 1.18
cuda/synchronization/stream/auto 714.4436090225564 ns 836.796875 ns 0.85
cuda/synchronization/stream/blocking 869.0555555555555 ns 660.0254777070064 ns 1.32
cuda/synchronization/stream/nonblocking 7479.5 ns 5573 ns 1.34
integration/byval/reference 147830 ns 147839 ns 1.00
integration/byval/slices=1 148976 ns 148934 ns 1.00
integration/byval/slices=2 291716 ns 291715 ns 1.00
integration/byval/slices=3 435055 ns 434876 ns 1.00
integration/cudadevrt 104984 ns 104869 ns 1.00
integration/volumerhs 9152686 ns 9148174 ns 1.00
kernel/indexing 12892 ns 12878 ns 1.00
kernel/indexing_checked 13754 ns 13404 ns 1.03
kernel/launch 2017.6666666666667 ns 2023.6666666666667 ns 1.00
kernel/occupancy 682.7615894039735 ns 691.2818791946308 ns 0.99
kernel/rand 17242 ns 16547 ns 1.04
latency/import 4238838358 ns 4229398317 ns 1.00
latency/precompile 4997986914 ns 4997574528 ns 1.00
latency/ttfp 4702000760 ns 4678764517 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

`used_memory` and `cached_memory` only need the device, but fetched it
through `active_state()`, which creates the task's default stream. They
are called by `maybe_collect` right before synchronizing, and creating a
stream can block the thread until the GPU is idle when the driver needs
to grow its pool of streams.
The default and legacy streams have no context, so handing them to a
worker thread failed. Pass the caller's context along instead. The
per-thread stream is specific to the calling thread, so wait for an
event recorded on it instead.

Also don't synchronize again on the calling thread when polling found
the object to have completed: that could block the thread on work that
another task submitted in the meantime. `isdone` already reports errors,
and a successful query counts as synchronization.
The tests relied on kernels running for a fixed amount of time, which
failed on CI when the thread was blocked creating a stream: the driver
waits for running kernels to finish when it grows its pool of streams.
Instead, keep the GPU busy until another task opens a gate, with a
timeout the tests check for, and set up all streams and kernels before.
Also cover the default, legacy and per-thread streams.
GPUToolbox 3.3.1 fixes two ways in which cooperative_wait's worker
threads could deadlock: a worker running a finalizer that waits for the
GPU (e.g. freeing memory) before notifying its waiter, and a wait for
an object that cannot be polled, like the context in device_synchronize,
waiting for a worker while all of them were busy.
@maleadt
maleadt added this pull request to stack #3319 October 1, 2026 11:24
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 86.10%. Comparing base (fde3437) to head (c05ddfa).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3315      +/-   ##
==========================================
+ Coverage   86.02%   86.10%   +0.08%     
==========================================
  Files         187      187              
  Lines       19042    18971      -71     
==========================================
- Hits        16380    16335      -45     
+ Misses       2662     2636      -26     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt merged commit a3202f2 into main Oct 2, 2026
1 check passed
@maleadt
maleadt deleted the tb/coopsync branch October 2, 2026 05:43
maleadt added a commit that referenced this pull request Oct 3, 2026
* Fix waiting for GPU work before releasing wrapped memory.

The release of wrapped system memory still called `nonblocking_synchronize`,
which #3315 removed. Waiting failed with an UndefVarError that was only logged,
after which the memory was unregistered and the Array released while the GPU
could still be using it. Use `cooperative_wait` instead.

Also wait and unregister with a relaxed capture mode: freeing a wrapper while
another stream is being captured, e.g. by the task doing the freeing, otherwise
invalidated that capture.

* Don't wait for a whole context while capturing.

Releasing wrapped memory last used on a special or destroyed stream waits for
the entire context, which is prohibited while any of its streams is being
captured, regardless of the capture mode. Postpone that until captures made
with `capture` have ended.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant