Skip to content

Hand off arrays between streams on the GPU instead of the host - #3329

Open
maleadt wants to merge 3 commits into
tb/unified_attachfrom
tb/device-handoff
Open

maleadt wants to merge 3 commits into
tb/unified_attachfrom
tb/device-handoff

Conversation

@maleadt

@maleadt maleadt commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

CUDA.jl remembers which stream last used each array. When the array is used on another stream, e.g., after passing it to another task, the previous stream is first synchronized from the host. That makes it safe to share arrays between tasks, but the task that takes over the array blocks until everything queued on the other stream has finished:

a = CUDA.rand(N)
long_running_kernel!(b)     # queued after the initialization of `a`

Threads.@spawn begin
    a .+= 1                 # blocks until `long_running_kernel!` is done
end

With this PR, operations that CUDA.jl submits itself (kernel launches, which includes broadcasts and KernelAbstractions kernels, graph launches, copies and fill!) make their stream wait for the previous one on the GPU instead, with cuStreamWaitEvent and an event that's cached per stream. The work still runs in the same order, but the host doesn't block.

Pointers that are passed to other code, like a library, MPI, or a ccall, may be used from the host or from other streams. Those conversions still synchronize, as before. That includes memory that was already handed off to the current stream: CUDA.jl remembers that the stream is waiting for another one, and synchronizes it when a pointer is taken. Pointers taken inside CUDA.with_managed are the exception: they're assumed to be used by an operation on the stream passed to with_managed, which is technically breaking for code that hands them to something else (noted in NEWS).

For this to be safe, an operation needs to hold the lock of the memory it uses until it has been submitted, so that another task can't see the new owner and wait for its stream before the work is there. Kernel launches already did that; the first commit makes copies and fill! do the same.

Some cases still synchronize on the host: unified and host memory, memory used by captures that CUDA.jl doesn't know about (i.e., not started with capture), the default, legacy and per-thread streams, and memory moving between devices. Library calls like cuBLAS could be handed off on the device too, but only after checking that each library queues all its work on the task's stream, so that's left for later.

Captures don't take ownership of the memory they use; launching the graph does. A graph launch is therefore handed off like any other operation: it waits on the device for the streams that last used its memory, wherever it was captured. A stream that is being captured with capture(; stream=s) needs care, because recording an event on it would only add a node to the graph. capture now records an event on s right before it begins capturing, under the same lock that hand-offs hold while recording on s, and other tasks wait for that event instead. That also lifts the documented limitation that other tasks couldn't use memory last used on s until the capture ended. The capturing task itself still gets a CaptureError when it waits for such memory, since the event doesn't cover what it captured.

On an RTX 5080, taking over an array while the other stream is busy for 0.7 s now returns after 0.1 ms instead of 0.7 s. Moving an array back and forth between two streams costs 2.4 µs per operation instead of 5.0 µs. Kernel launches cost the same as before. Copies and fill! take about 50–70 ns longer (~1.9 µs and ~0.9 µs before), because the arrays now stay locked until the operation has been submitted, like kernel launches already did.

The new tests keep a stream busy with a gate kernel. They check that kernels and copies in other tasks don't block but still execute in order, that synchronize(a) waits for the stream the array came from, that a pointer conversion after a hand-off still waits for the original stream, that unified memory still synchronizes, and that a graph captured right after a hand-off and launched on another stream doesn't block and sees the right data. The graph tests check that other tasks can use memory last used on a stream while it's being captured on explicitly, and that the capture itself can't. Tested with and without compute-sanitizer, and on Julia 1.11 for the launch-allocation test.

@maleadt
maleadt added this pull request to stack #3330 October 4, 2026 07:08
@maleadt maleadt changed the title Optionally order memory hand-offs between streams on the device Optionally hand off memory between streams on the device Oct 4, 2026
@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.73684% with 5 lines in your changes missing coverage. Please review.
✅ Project coverage is 40.78%. Comparing base (a2f881a) to head (7f41843).

Files with missing lines Patch % Lines
CUDACore/src/resources/managed.jl 88.46% 3 Missing ⚠️
CUDACore/src/array.jl 92.59% 2 Missing ⚠️
Additional details and impacted files
@@                  Coverage Diff                  @@
##           tb/unified_attach    #3329      +/-   ##
=====================================================
- Coverage              41.38%   40.78%   -0.61%     
=====================================================
  Files                    194      194              
  Lines                  19686    19249     -437     
=====================================================
- Hits                    8148     7851     -297     
+ Misses                 11538    11398     -140     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt removed this pull request from stack #3330 October 4, 2026 08:26
@maleadt
maleadt changed the base branch from tb/event-ownership to main October 4, 2026 08:26
@maleadt
maleadt force-pushed the tb/device-handoff branch from f4f596b to f4cfe15 Compare October 4, 2026 08:26
@maleadt
maleadt added this pull request to stack #3333 October 4, 2026 08:26
@maleadt maleadt changed the title Optionally hand off memory between streams on the device Hand off arrays between streams on the GPU instead of the host Oct 4, 2026
@maleadt
maleadt removed this pull request from stack #3333 October 4, 2026 09:03
@maleadt
maleadt force-pushed the tb/device-handoff branch from f4cfe15 to 0944b01 Compare October 4, 2026 09:03
@maleadt
maleadt changed the base branch from main to tb/deferred-release October 4, 2026 09:04
@maleadt
maleadt added this pull request to stack #3334 October 4, 2026 09:04
@maleadt
maleadt force-pushed the tb/device-handoff branch from 0944b01 to 5ee5e5b Compare October 4, 2026 09:04
@maleadt
maleadt force-pushed the tb/device-handoff branch from 5ee5e5b to 7f41843 Compare October 5, 2026 08:22
@maleadt
maleadt removed this pull request from stack #3334 October 5, 2026 08:33
@maleadt
maleadt changed the base branch from tb/deferred-release to tb/unified_attach October 5, 2026 08:33
@maleadt
maleadt added this pull request to stack #3338 October 5, 2026 08:33
@maleadt
maleadt marked this pull request as ready for review October 5, 2026 09:08
@maleadt
maleadt force-pushed the tb/device-handoff branch from 7f41843 to 75b8119 Compare October 5, 2026 18:16
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 10b60ad Previous: e924726 Ratio
array/accumulate/Float32/1d 99991 ns 99107 ns 1.01
array/accumulate/Float32/dims=1 74340 ns 72803 ns 1.02
array/accumulate/Float32/dims=1L 1589434 ns 1589176 ns 1.00
array/accumulate/Float32/dims=2 139991 ns 138542 ns 1.01
array/accumulate/Float32/dims=2L 658383 ns 655298 ns 1.00
array/accumulate/Int64/1d 119136 ns 117327 ns 1.02
array/accumulate/Int64/dims=1 78789 ns 77135 ns 1.02
array/accumulate/Int64/dims=1L 1704174 ns 1698128 ns 1.00
array/accumulate/Int64/dims=2 154641 ns 150131 ns 1.03
array/accumulate/Int64/dims=2L 989615 ns 986795 ns 1.00
array/broadcast 16269 ns 16029 ns 1.01
array/broadcast launch 7510.75 ns 7358.25 ns 1.02
array/construct 981.4 ns 959.9 ns 1.02
array/copy 17323 ns 16576 ns 1.05
array/copyto!/cpu_to_gpu 210396 ns 207398 ns 1.01
array/copyto!/gpu_to_cpu 242214 ns 239800 ns 1.01
array/copyto!/gpu_to_gpu 8911.666666666666 ns 8720.666666666666 ns 1.02
array/iteration/findall/bool 135615 ns 129599 ns 1.05
array/iteration/findall/int 145085 ns 140825 ns 1.03
array/iteration/findfirst/bool 73026 ns 67856 ns 1.08
array/iteration/findfirst/int 75000 ns 69203 ns 1.08
array/iteration/findmin/1d 69282 ns 59849 ns 1.16
array/iteration/findmin/2d 97666 ns 97040 ns 1.01
array/iteration/logical 196471 ns 181916 ns 1.08
array/iteration/scalar 57054 ns 58365 ns 0.98
array/permutedims/2d 46976 ns 45704 ns 1.03
array/permutedims/3d 48647 ns 47728 ns 1.02
array/permutedims/4d 49385 ns 48656 ns 1.01
array/random/rand/Float32 11828 ns 10851 ns 1.09
array/random/rand/Int64 19499 ns 18300 ns 1.07
array/random/rand!/Float32 7751 ns 7842.666666666667 ns 0.99
array/random/rand!/Int64 17045 ns 16741 ns 1.02
array/random/randn/Float32 33102 ns 32264 ns 1.03
array/random/randn!/Float32 24394 ns 24124 ns 1.01
array/reductions/mapreduce/Float32/1d 36648 ns 32799 ns 1.12
array/reductions/mapreduce/Float32/dims=1 38785 ns 37628 ns 1.03
array/reductions/mapreduce/Float32/dims=1L 52170 ns 51393 ns 1.02
array/reductions/mapreduce/Float32/dims=2 56400 ns 55676 ns 1.01
array/reductions/mapreduce/Float32/dims=2L 68833 ns 67890 ns 1.01
array/reductions/mapreduce/Int64/1d 43320 ns 39071 ns 1.11
array/reductions/mapreduce/Int64/dims=1 40877 ns 40979 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 89492 ns 89303 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57595 ns 57656 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 85312 ns 84117 ns 1.01
array/reductions/reduce/Float32/1d 36980 ns 32726 ns 1.13
array/reductions/reduce/Float32/dims=1 38576 ns 37671 ns 1.02
array/reductions/reduce/Float32/dims=1L 51736 ns 51035 ns 1.01
array/reductions/reduce/Float32/dims=2 56243 ns 55445 ns 1.01
array/reductions/reduce/Float32/dims=2L 69370 ns 68151 ns 1.02
array/reductions/reduce/Int64/1d 43256 ns 39207 ns 1.10
array/reductions/reduce/Int64/dims=1 41125 ns 40842 ns 1.01
array/reductions/reduce/Int64/dims=1L 89615 ns 88981 ns 1.01
array/reductions/reduce/Int64/dims=2 57467 ns 58095 ns 0.99
array/reductions/reduce/Int64/dims=2L 84821 ns 84484 ns 1.00
array/reverse/1d 17792 ns 17482 ns 1.02
array/reverse/1dL 70481 ns 70086 ns 1.01
array/reverse/1dL_inplace 67889 ns 67824 ns 1.00
array/reverse/1d_inplace 9046 ns 9051.333333333334 ns 1.00
array/reverse/2d 21064 ns 20350 ns 1.04
array/reverse/2dL 74902 ns 73763 ns 1.02
array/reverse/2dL_inplace 67545 ns 67393 ns 1.00
array/reverse/2d_inplace 12508 ns 10090 ns 1.24
array/sorting/1d 2645238 ns 2646314 ns 1.00
array/sorting/2d 1019579 ns 1018011 ns 1.00
array/sorting/by 3175666 ns 3158564 ns 1.01
cuda/graph/capture 45373 ns
cuda/graph/eager 39596 ns
cuda/graph/launch 13438 ns
cuda/synchronization/context/auto 8025.333333333333 ns 6775.2 ns 1.18
cuda/synchronization/context/blocking 875.6086956521739 ns 808.9775280898876 ns 1.08
cuda/synchronization/context/nonblocking 7921 ns 6745.2 ns 1.17
cuda/synchronization/stream/auto 737.511811023622 ns 707.3802816901408 ns 1.04
cuda/synchronization/stream/blocking 949.5652173913044 ns 862.0714285714286 ns 1.10
cuda/synchronization/stream/nonblocking 8684.666666666666 ns 7079.5 ns 1.23
integration/byval/reference 148763 ns 148360 ns 1.00
integration/byval/slices=1 149783 ns 149203 ns 1.00
integration/byval/slices=2 292424 ns 292017 ns 1.00
integration/byval/slices=3 435524 ns 434881 ns 1.00
integration/cudadevrt 105490 ns 105387 ns 1.00
integration/volumerhs 9150566 ns 9145562 ns 1.00
kernel/indexing 13631 ns 13190 ns 1.03
kernel/indexing_checked 14425 ns 13943 ns 1.03
kernel/launch 2492.1111111111113 ns 2472 ns 1.01
kernel/occupancy 957.05 ns 938.1739130434783 ns 1.02
kernel/rand 15235 ns 14010 ns 1.09
latency/import 4317343074 ns 4302362850 ns 1.00
latency/precompile 5107515720 ns 5082487750 ns 1.00
latency/ttfp 4822706938 ns 4807353177 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Kernel launches keep the memory they use locked until the launch has
been submitted, so that another task can't see the new owner of the
memory and wait for its stream before the work it needs to wait for is
on there. Copies and fill! took ownership when converting their arrays
to pointers, but released the lock again before submitting the
operation. Keep their arrays locked until the operation has been
submitted too.
When an array is used on another stream than the one that last used it,
e.g., because it was passed to another task, CUDA.jl synchronizes the
previous stream from the host. That makes it safe to share arrays
between tasks, but the task that takes over the array blocks until all
work on the other stream has finished, including work that has nothing
to do with the array.

For operations that CUDA.jl submits itself (kernel launches, graph
launches, copies and fills), make the new stream wait for the previous
one on the device instead, using an event that's cached per stream. The
operations still execute in order, but the host isn't blocked. Pointers
to memory that are taken outside of such an operation, e.g., by a
library or MPI, may be used from the host or other streams, so these
conversions still synchronize on the host. That also applies when the
memory was already handed off to the current stream: the memory
remembers that its stream waits for others, and a pointer conversion
then synchronizes that stream.

Only device memory moving between ordinary streams in the same context
is handed off on the device. Unified and host memory, memory used by
captures that CUDA.jl doesn't know about, and the special streams keep
using host synchronization. So does a stream that is being captured on
explicitly, as recording an event on it would only add a node to the
graph; `capture` marks the stream under the lock that hand-offs hold
while recording, before it begins capturing, so that no hand-off can
record an event on it after that. Captures don't take ownership of the
memory they use, launching the graph does, so a graph launch is ordered
after earlier uses of its memory like any other operation.
When capturing on a specific stream, memory may have been used on that
stream before the capture. Another task that wants to use that memory
needs to wait for the work that was submitted to the stream, but can't
do so by recording an event on it: during the capture, that would only
add a node to the graph. So such uses failed with a CaptureError until
the capture had ended.

Record an event on the stream right before beginning the capture, and
have other tasks wait for that event instead, when handing memory off on
the device, when synchronizing it from the host, and when attaching
unified memory to the host before CPU access (which then happens on
another stream). Captured operations don't take ownership of memory, so
the event covers the last use of all memory owned by the stream. It is
published under the lock that is held while recording events on the
stream, so that no event is recorded on it after the capture began. Each
capture records a new event, as tasks may still be waiting for the one
of an earlier capture. The capture itself keeps failing to synchronize
such memory, as it may have captured operations on it that the event
doesn't cover.
@maleadt
maleadt force-pushed the tb/device-handoff branch from 10b60ad to 744ca5a Compare October 7, 2026 04:54

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant