Repository navigation
Conversation
maleadt
added this pull request to stack #3330
October 4, 2026 07:08
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/unified_attach #3329 +/- ##
=====================================================
- Coverage 41.38% 40.78% -0.61%
=====================================================
Files 194 194
Lines 19686 19249 -437
=====================================================
- Hits 8148 7851 -297
+ Misses 11538 11398 -140 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
maleadt
removed this pull request from stack #3330
October 4, 2026 08:26
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 4, 2026 08:26
f4f596b to
f4cfe15
Compare
maleadt
added this pull request to stack #3333
October 4, 2026 08:26
maleadt
removed this pull request from stack #3333
October 4, 2026 09:03
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 4, 2026 09:03
f4cfe15 to
0944b01
Compare
maleadt
added this pull request to stack #3334
October 4, 2026 09:04
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 4, 2026 09:04
0944b01 to
5ee5e5b
Compare
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 5, 2026 08:22
5ee5e5b to
7f41843
Compare
maleadt
removed this pull request from stack #3334
October 5, 2026 08:33
maleadt
added this pull request to stack #3338
October 5, 2026 08:33
maleadt
marked this pull request as ready for review
October 5, 2026 09:08
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 5, 2026 18:16
7f41843 to
75b8119
Compare
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 6, 2026 08:24
75b8119 to
de8639e
Compare
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 6, 2026 13:55
de8639e to
10b60ad
Compare
Kernel launches keep the memory they use locked until the launch has been submitted, so that another task can't see the new owner of the memory and wait for its stream before the work it needs to wait for is on there. Copies and fill! took ownership when converting their arrays to pointers, but released the lock again before submitting the operation. Keep their arrays locked until the operation has been submitted too.
When an array is used on another stream than the one that last used it, e.g., because it was passed to another task, CUDA.jl synchronizes the previous stream from the host. That makes it safe to share arrays between tasks, but the task that takes over the array blocks until all work on the other stream has finished, including work that has nothing to do with the array. For operations that CUDA.jl submits itself (kernel launches, graph launches, copies and fills), make the new stream wait for the previous one on the device instead, using an event that's cached per stream. The operations still execute in order, but the host isn't blocked. Pointers to memory that are taken outside of such an operation, e.g., by a library or MPI, may be used from the host or other streams, so these conversions still synchronize on the host. That also applies when the memory was already handed off to the current stream: the memory remembers that its stream waits for others, and a pointer conversion then synchronizes that stream. Only device memory moving between ordinary streams in the same context is handed off on the device. Unified and host memory, memory used by captures that CUDA.jl doesn't know about, and the special streams keep using host synchronization. So does a stream that is being captured on explicitly, as recording an event on it would only add a node to the graph; `capture` marks the stream under the lock that hand-offs hold while recording, before it begins capturing, so that no hand-off can record an event on it after that. Captures don't take ownership of the memory they use, launching the graph does, so a graph launch is ordered after earlier uses of its memory like any other operation.
When capturing on a specific stream, memory may have been used on that stream before the capture. Another task that wants to use that memory needs to wait for the work that was submitted to the stream, but can't do so by recording an event on it: during the capture, that would only add a node to the graph. So such uses failed with a CaptureError until the capture had ended. Record an event on the stream right before beginning the capture, and have other tasks wait for that event instead, when handing memory off on the device, when synchronizing it from the host, and when attaching unified memory to the host before CPU access (which then happens on another stream). Captured operations don't take ownership of memory, so the event covers the last use of all memory owned by the stream. It is published under the lock that is held while recording events on the stream, so that no event is recorded on it after the capture began. Each capture records a new event, as tasks may still be waiting for the one of an earlier capture. The capture itself keeps failing to synchronize such memory, as it may have captured operations on it that the event doesn't cover.
maleadt
force-pushed
the
tb/device-handoff
branch
from
October 7, 2026 04:54
10b60ad to
744ca5a
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CUDA.jl remembers which stream last used each array. When the array is used on another stream, e.g., after passing it to another task, the previous stream is first synchronized from the host. That makes it safe to share arrays between tasks, but the task that takes over the array blocks until everything queued on the other stream has finished:
With this PR, operations that CUDA.jl submits itself (kernel launches, which includes broadcasts and KernelAbstractions kernels, graph launches, copies and
fill!) make their stream wait for the previous one on the GPU instead, withcuStreamWaitEventand an event that's cached per stream. The work still runs in the same order, but the host doesn't block.Pointers that are passed to other code, like a library, MPI, or a
ccall, may be used from the host or from other streams. Those conversions still synchronize, as before. That includes memory that was already handed off to the current stream: CUDA.jl remembers that the stream is waiting for another one, and synchronizes it when a pointer is taken. Pointers taken insideCUDA.with_managedare the exception: they're assumed to be used by an operation on the stream passed towith_managed, which is technically breaking for code that hands them to something else (noted in NEWS).For this to be safe, an operation needs to hold the lock of the memory it uses until it has been submitted, so that another task can't see the new owner and wait for its stream before the work is there. Kernel launches already did that; the first commit makes copies and
fill!do the same.Some cases still synchronize on the host: unified and host memory, memory used by captures that CUDA.jl doesn't know about (i.e., not started with
capture), the default, legacy and per-thread streams, and memory moving between devices. Library calls like cuBLAS could be handed off on the device too, but only after checking that each library queues all its work on the task's stream, so that's left for later.Captures don't take ownership of the memory they use; launching the graph does. A graph launch is therefore handed off like any other operation: it waits on the device for the streams that last used its memory, wherever it was captured. A stream that is being captured with
capture(; stream=s)needs care, because recording an event on it would only add a node to the graph.capturenow records an event onsright before it begins capturing, under the same lock that hand-offs hold while recording ons, and other tasks wait for that event instead. That also lifts the documented limitation that other tasks couldn't use memory last used onsuntil the capture ended. The capturing task itself still gets aCaptureErrorwhen it waits for such memory, since the event doesn't cover what it captured.On an RTX 5080, taking over an array while the other stream is busy for 0.7 s now returns after 0.1 ms instead of 0.7 s. Moving an array back and forth between two streams costs 2.4 µs per operation instead of 5.0 µs. Kernel launches cost the same as before. Copies and
fill!take about 50–70 ns longer (~1.9 µs and ~0.9 µs before), because the arrays now stay locked until the operation has been submitted, like kernel launches already did.The new tests keep a stream busy with a gate kernel. They check that kernels and copies in other tasks don't block but still execute in order, that
synchronize(a)waits for the stream the array came from, that a pointer conversion after a hand-off still waits for the original stream, that unified memory still synchronizes, and that a graph captured right after a hand-off and launched on another stream doesn't block and sees the right data. The graph tests check that other tasks can use memory last used on a stream while it's being captured on explicitly, and that the capture itself can't. Tested with and without compute-sanitizer, and on Julia 1.11 for the launch-allocation test.