Repository navigation
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/device-handoff #3328 +/- ##
=====================================================
+ Coverage 40.78% 41.08% +0.29%
=====================================================
Files 194 194
Lines 19249 19350 +101
=====================================================
+ Hits 7851 7950 +99
- Misses 11398 11400 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 4, 2026 06:29
8c291ff to
d129f2c
Compare
maleadt
added this pull request to stack #3330
October 4, 2026 07:08
maleadt
removed this pull request from stack #3330
October 4, 2026 08:26
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 4, 2026 08:26
d129f2c to
c1db358
Compare
maleadt
added this pull request to stack #3333
October 4, 2026 08:26
maleadt
removed this pull request from stack #3333
October 4, 2026 09:03
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 4, 2026 09:04
c1db358 to
e2056f6
Compare
maleadt
added this pull request to stack #3334
October 4, 2026 09:04
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 4, 2026 09:04
e2056f6 to
a80c979
Compare
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 5, 2026 08:22
a80c979 to
598e849
Compare
maleadt
removed this pull request from stack #3334
October 5, 2026 08:33
maleadt
added this pull request to stack #3338
October 5, 2026 08:33
maleadt
marked this pull request as ready for review
October 5, 2026 09:08
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 5, 2026 18:16
598e849 to
0343fe7
Compare
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
maleadt
force-pushed
the
tb/event-ownership
branch
2 times, most recently
from
October 6, 2026 13:55
7b06014 to
636c841
Compare
When memory moves to another stream, CUDA.jl orders the new stream after the one that last used the memory, by waiting on the device or synchronizing on the host. That's unnecessary when the host already knows that the last use has completed, e.g., because it synchronized an event recorded after it. Divide each stream's work into epochs, stamp memory with the epoch of its last use, and close an epoch when recording an event or synchronizing the stream. Synchronizing a stream, or observing that an event has completed, then tells us that all accesses up to the closed epoch have finished, so memory last used in them doesn't need to be waited for anymore. Waiting for an event doesn't hold its lock, so other tasks can record or query it in the meantime; what the wait covers is only learned if the event wasn't recorded again. Operations that CUDA.jl submits itself (kernel launches, copies, fill!, graph launches) stamp memory when taking its pointer, and again once the operation has been submitted, also when that fails. Between the two, an epoch closed by another task could appear to cover the operation, so only the task submitting work to a stream closes epochs that count. Epochs closed by another task, or on a stream that several live tasks submit work to, therefore don't cover anything. A stream whose task has finished can be taken over by another one, as happens when streams are recycled. A pointer taken outside of such an operation, e.g. with `pointer(a)` or when passing an array to a library call, is only stamped once and can be used by work that is submitted after its epoch was closed. Such memory isn't considered complete until it moves to another stream, which waits for all work submitted to the old one so far. Recording an event on a stream that is being captured only adds a node to the graph, so it doesn't close an epoch, and captured operations don't stamp the memory they use (launching the graph does). Memory used by captures that CUDA.jl doesn't know about, and memory used on the special streams, are never considered complete.
When a stream waits for an event, everything it executes afterwards is ordered after the work the event covers. Memory last used in that work can then move to that stream without waiting for the stream that used it. That's what KernelAbstractions' `@spawn` relies on: the spawned task waits for an event recorded by its parent, but its first use of the parent's arrays currently also waits for everything the parent queued after the event. Have each stream remember, per stream it waited for, up to which epoch its work is ordered, and don't wait again for memory whose last use is covered by that. This applies to the memory that could be handed off on the device anyway, and not to memory whose pointer escaped. The memory remains marked as waiting on another stream, so that a pointer conversion for a consumer outside of CUDA.jl, like MPI, synchronizes the new stream. That only waits for the work the event covered, and not for everything the other stream queued since, as synchronizing that stream would.
maleadt
force-pushed
the
tb/event-ownership
branch
from
October 7, 2026 04:54
636c841 to
5903e78
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When memory moves between streams, CUDA.jl makes the new stream wait for everything the previous stream had queued. That's more than needed when the program already ordered the work with an event.
KernelAbstractions.@spawnis the typical example (JuliaGPU/KernelAbstractions.jl#817): the parent records an event and the child's stream waits for it. Still, the child's first use of an array waited for all the work the parent queued after the event, so the two tasks didn't overlap:To know what an event covers, each stream's work is split into epochs. Memory is stamped with the epoch of its last use, and recording an event or synchronizing a stream closes the current epoch. That gives two facts that make waiting unnecessary:
isdone(event)returned true.CUDA.wait(event, stream). Work submitted to it afterwards runs after everything the event covers.Memory covered by either fact moves to the new stream without waiting. A pointer conversion for code outside CUDA.jl, like the MPI call above, then synchronizes the new stream instead of the parent's, which only waits for what the event covered. Oceananigans can therefore drop its workaround of making the task adopt the parent's stream (CliMA/Oceananigans.jl#5897).
Operations that CUDA.jl submits (kernel launches, copies,
fill!, graph launches) stamp memory when they take its pointer and again once the operation has been submitted. Between the two, an epoch closed by another task could appear to cover the operation, so an epoch only counts if the task that submits work to the stream closed it. Events recorded on another task's stream, or on a stream that several live tasks use, cover nothing. A stream whose task has finished can be taken over by another task, which recycling streams relies on. Graph capture doesn't produce either fact: recording an event on a stream that is being captured only adds a node to the graph, a wait on a capturing stream only orders the graph's operations, and captured operations don't stamp memory (launching the graph does).A pointer taken outside of a CUDA.jl operation, e.g.
p = pointer(a)passed to a kernel or library later, or an array passed to a library call, is only stamped once and can be used after the epoch it was stamped with has been closed:p = pointer(a); synchronize(); @cuda f(p)isn't covered by that synchronization. Such memory doesn't use either shortcut until it moves to another stream, which waits for everything submitted to the old one, as before. Work submitted through the old pointer after the array was used by another task isn't ordered with that task's work; that was already the case, and is now documented. Waiting for an event doesn't hold its lock, so other tasks can record or query it meanwhile; a wait only counts if the event wasn't recorded again in the meantime.On an RTX 5080,
recordandCUDA.waittake about 150 and 210 ns, up from about 100 ns each, because of a lock and a table of the streams that were waited for. Kernel launches and copies cost the same as before.The tests use pointer conversions and stream synchronization to check that events synchronized on the host, events waited for by a stream, and streams taken over from finished tasks avoid waiting. They also check that memory used after an event, events recorded by another task, events re-recorded elsewhere (also while another task waits for them), and launches through a pointer taken before the event or synchronization still wait.