Skip to content

Don't wait for work that an event already ordered - #3328

Open
maleadt wants to merge 2 commits into
tb/device-handofffrom
tb/event-ownership
Open

maleadt wants to merge 2 commits into
tb/device-handofffrom
tb/event-ownership

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

When memory moves between streams, CUDA.jl makes the new stream wait for everything the previous stream had queued. That's more than needed when the program already ordered the work with an event. KernelAbstractions.@spawn is the typical example (JuliaGPU/KernelAbstractions.jl#817): the parent records an event and the child's stream waits for it. Still, the child's first use of an array waited for all the work the parent queued after the event, so the two tasks didn't overlap:

a .= 1
e = CuEvent(); record(e)
more_work!(b)                   # the parent keeps going

Threads.@spawn begin            # what KernelAbstractions.@spawn does
    CUDA.wait(e)
    a .+= 1                     # doesn't wait for `more_work!` anymore
    MPI.Send(a, comm; dest=1)   # waits for `e`, but not for `more_work!`
end

To know what an event covers, each stream's work is split into epochs. Memory is stamped with the epoch of its last use, and recording an event or synchronizing a stream closes the current epoch. That gives two facts that make waiting unnecessary:

  • Completed: the host saw the work finish, by synchronizing the stream or the event, or because isdone(event) returned true.
  • Ordered: a stream waited for an event with CUDA.wait(event, stream). Work submitted to it afterwards runs after everything the event covers.

Memory covered by either fact moves to the new stream without waiting. A pointer conversion for code outside CUDA.jl, like the MPI call above, then synchronizes the new stream instead of the parent's, which only waits for what the event covered. Oceananigans can therefore drop its workaround of making the task adopt the parent's stream (CliMA/Oceananigans.jl#5897).

Operations that CUDA.jl submits (kernel launches, copies, fill!, graph launches) stamp memory when they take its pointer and again once the operation has been submitted. Between the two, an epoch closed by another task could appear to cover the operation, so an epoch only counts if the task that submits work to the stream closed it. Events recorded on another task's stream, or on a stream that several live tasks use, cover nothing. A stream whose task has finished can be taken over by another task, which recycling streams relies on. Graph capture doesn't produce either fact: recording an event on a stream that is being captured only adds a node to the graph, a wait on a capturing stream only orders the graph's operations, and captured operations don't stamp memory (launching the graph does).

A pointer taken outside of a CUDA.jl operation, e.g. p = pointer(a) passed to a kernel or library later, or an array passed to a library call, is only stamped once and can be used after the epoch it was stamped with has been closed: p = pointer(a); synchronize(); @cuda f(p) isn't covered by that synchronization. Such memory doesn't use either shortcut until it moves to another stream, which waits for everything submitted to the old one, as before. Work submitted through the old pointer after the array was used by another task isn't ordered with that task's work; that was already the case, and is now documented. Waiting for an event doesn't hold its lock, so other tasks can record or query it meanwhile; a wait only counts if the event wasn't recorded again in the meantime.

On an RTX 5080, record and CUDA.wait take about 150 and 210 ns, up from about 100 ns each, because of a lock and a table of the streams that were waited for. Kernel launches and copies cost the same as before.

The tests use pointer conversions and stream synchronization to check that events synchronized on the host, events waited for by a stream, and streams taken over from finished tasks avoid waiting. They also check that memory used after an event, events recorded by another task, events re-recorded elsewhere (also while another task waits for them), and launches through a pointer taken before the event or synchronization still wait.

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.18033% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 41.08%. Comparing base (7f41843) to head (598e849).

Files with missing lines Patch % Lines
CUDACore/lib/cudadrv/stream.jl 96.42% 1 Missing ⚠️
Additional details and impacted files
@@                  Coverage Diff                  @@
##           tb/device-handoff    #3328      +/-   ##
=====================================================
+ Coverage              40.78%   41.08%   +0.29%     
=====================================================
  Files                    194      194              
  Lines                  19249    19350     +101     
=====================================================
+ Hits                    7851     7950      +99     
- Misses                 11398    11400       +2     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt force-pushed the tb/event-ownership branch from 8c291ff to d129f2c Compare October 4, 2026 06:29
@maleadt
maleadt added this pull request to stack #3330 October 4, 2026 07:08
@maleadt
maleadt removed this pull request from stack #3330 October 4, 2026 08:26
@maleadt
maleadt force-pushed the tb/event-ownership branch from d129f2c to c1db358 Compare October 4, 2026 08:26
@maleadt
maleadt changed the base branch from main to tb/device-handoff October 4, 2026 08:26
@maleadt
maleadt added this pull request to stack #3333 October 4, 2026 08:26
@maleadt maleadt changed the title Don't synchronize memory that's already ordered by an event Don't wait for work that an event already ordered Oct 4, 2026
@maleadt
maleadt removed this pull request from stack #3333 October 4, 2026 09:03
@maleadt
maleadt force-pushed the tb/event-ownership branch from c1db358 to e2056f6 Compare October 4, 2026 09:04
@maleadt
maleadt added this pull request to stack #3334 October 4, 2026 09:04
@maleadt
maleadt force-pushed the tb/event-ownership branch from e2056f6 to a80c979 Compare October 4, 2026 09:04
@maleadt
maleadt force-pushed the tb/event-ownership branch from a80c979 to 598e849 Compare October 5, 2026 08:22
@maleadt
maleadt removed this pull request from stack #3334 October 5, 2026 08:33
@maleadt
maleadt added this pull request to stack #3338 October 5, 2026 08:33
@maleadt
maleadt marked this pull request as ready for review October 5, 2026 09:08
@maleadt
maleadt force-pushed the tb/event-ownership branch from 598e849 to 0343fe7 Compare October 5, 2026 18:16
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 636c841 Previous: e924726 Ratio
array/accumulate/Float32/1d 99015 ns 99107 ns 1.00
array/accumulate/Float32/dims=1 73513 ns 72803 ns 1.01
array/accumulate/Float32/dims=1L 1588632 ns 1589176 ns 1.00
array/accumulate/Float32/dims=2 139356 ns 138542 ns 1.01
array/accumulate/Float32/dims=2L 656715 ns 655298 ns 1.00
array/accumulate/Int64/1d 118052 ns 117327 ns 1.01
array/accumulate/Int64/dims=1 78031 ns 77135 ns 1.01
array/accumulate/Int64/dims=1L 1702573 ns 1698128 ns 1.00
array/accumulate/Int64/dims=2 153684 ns 150131 ns 1.02
array/accumulate/Int64/dims=2L 988212 ns 986795 ns 1.00
array/broadcast 16367 ns 16029 ns 1.02
array/broadcast launch 7637.75 ns 7358.25 ns 1.04
array/construct 958.125 ns 959.9 ns 1.00
array/copy 17200 ns 16576 ns 1.04
array/copyto!/cpu_to_gpu 207298 ns 207398 ns 1.00
array/copyto!/gpu_to_cpu 240367 ns 239800 ns 1.00
array/copyto!/gpu_to_gpu 8939 ns 8720.666666666666 ns 1.03
array/iteration/findall/bool 134512 ns 129599 ns 1.04
array/iteration/findall/int 145316 ns 140825 ns 1.03
array/iteration/findfirst/bool 71734 ns 67856 ns 1.06
array/iteration/findfirst/int 73900 ns 69203 ns 1.07
array/iteration/findmin/1d 71448 ns 59849 ns 1.19
array/iteration/findmin/2d 97774 ns 97040 ns 1.01
array/iteration/logical 196056 ns 181916 ns 1.08
array/iteration/scalar 58401 ns 58365 ns 1.00
array/permutedims/2d 46855 ns 45704 ns 1.03
array/permutedims/3d 48249 ns 47728 ns 1.01
array/permutedims/4d 49322 ns 48656 ns 1.01
array/random/rand/Float32 11672 ns 10851 ns 1.08
array/random/rand/Int64 19599 ns 18300 ns 1.07
array/random/rand!/Float32 7754.333333333333 ns 7842.666666666667 ns 0.99
array/random/rand!/Int64 17075 ns 16741 ns 1.02
array/random/randn/Float32 32740 ns 32264 ns 1.01
array/random/randn!/Float32 23531 ns 24124 ns 0.98
array/reductions/mapreduce/Float32/1d 35357 ns 32799 ns 1.08
array/reductions/mapreduce/Float32/dims=1 38008 ns 37628 ns 1.01
array/reductions/mapreduce/Float32/dims=1L 52077 ns 51393 ns 1.01
array/reductions/mapreduce/Float32/dims=2 55631 ns 55676 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 68227 ns 67890 ns 1.00
array/reductions/mapreduce/Int64/1d 42435 ns 39071 ns 1.09
array/reductions/mapreduce/Int64/dims=1 41065 ns 40979 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 89570 ns 89303 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57340 ns 57656 ns 0.99
array/reductions/mapreduce/Int64/dims=2L 85226 ns 84117 ns 1.01
array/reductions/reduce/Float32/1d 35718 ns 32726 ns 1.09
array/reductions/reduce/Float32/dims=1 38088 ns 37671 ns 1.01
array/reductions/reduce/Float32/dims=1L 51707 ns 51035 ns 1.01
array/reductions/reduce/Float32/dims=2 55822 ns 55445 ns 1.01
array/reductions/reduce/Float32/dims=2L 68980 ns 68151 ns 1.01
array/reductions/reduce/Int64/1d 41755 ns 39207 ns 1.06
array/reductions/reduce/Int64/dims=1 41178 ns 40842 ns 1.01
array/reductions/reduce/Int64/dims=1L 89529 ns 88981 ns 1.01
array/reductions/reduce/Int64/dims=2 57315 ns 58095 ns 0.99
array/reductions/reduce/Int64/dims=2L 84974 ns 84484 ns 1.01
array/reverse/1d 17386 ns 17482 ns 0.99
array/reverse/1dL 70574 ns 70086 ns 1.01
array/reverse/1dL_inplace 67842 ns 67824 ns 1.00
array/reverse/1d_inplace 10771 ns 9051.333333333334 ns 1.19
array/reverse/2d 20650 ns 20350 ns 1.01
array/reverse/2dL 74438 ns 73763 ns 1.01
array/reverse/2dL_inplace 67596 ns 67393 ns 1.00
array/reverse/2d_inplace 12529 ns 10090 ns 1.24
array/sorting/1d 2647055 ns 2646314 ns 1.00
array/sorting/2d 1018744 ns 1018011 ns 1.00
array/sorting/by 3174481 ns 3158564 ns 1.01
cuda/graph/capture 42934 ns
cuda/graph/eager 42703 ns
cuda/graph/launch 15390 ns
cuda/synchronization/context/auto 7832.666666666667 ns 6775.2 ns 1.16
cuda/synchronization/context/blocking 874.3269230769231 ns 808.9775280898876 ns 1.08
cuda/synchronization/context/nonblocking 7842.5 ns 6745.2 ns 1.16
cuda/synchronization/stream/auto 729.3587786259542 ns 707.3802816901408 ns 1.03
cuda/synchronization/stream/blocking 954.88 ns 862.0714285714286 ns 1.11
cuda/synchronization/stream/nonblocking 8198.333333333334 ns 7079.5 ns 1.16
integration/byval/reference 148491 ns 148360 ns 1.00
integration/byval/slices=1 149652 ns 149203 ns 1.00
integration/byval/slices=2 292130 ns 292017 ns 1.00
integration/byval/slices=3 435207 ns 434881 ns 1.00
integration/cudadevrt 105662 ns 105387 ns 1.00
integration/volumerhs 9155562 ns 9145562 ns 1.00
kernel/indexing 13529 ns 13190 ns 1.03
kernel/indexing_checked 14113 ns 13943 ns 1.01
kernel/launch 2503.777777777778 ns 2472 ns 1.01
kernel/occupancy 990.1333333333333 ns 938.1739130434783 ns 1.06
kernel/rand 16100 ns 14010 ns 1.15
latency/import 4327222830 ns 4302362850 ns 1.01
latency/precompile 5095124846 ns 5082487750 ns 1.00
latency/ttfp 4851067301 ns 4807353177 ns 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@maleadt
maleadt force-pushed the tb/event-ownership branch 2 times, most recently from 7b06014 to 636c841 Compare October 6, 2026 13:55
When memory moves to another stream, CUDA.jl orders the new stream after
the one that last used the memory, by waiting on the device or
synchronizing on the host. That's unnecessary when the host already knows
that the last use has completed, e.g., because it synchronized an event
recorded after it.

Divide each stream's work into epochs, stamp memory with the epoch of its
last use, and close an epoch when recording an event or synchronizing the
stream. Synchronizing a stream, or observing that an event has completed,
then tells us that all accesses up to the closed epoch have finished, so
memory last used in them doesn't need to be waited for anymore. Waiting
for an event doesn't hold its lock, so other tasks can record or query it
in the meantime; what the wait covers is only learned if the event wasn't
recorded again.

Operations that CUDA.jl submits itself (kernel launches, copies, fill!,
graph launches) stamp memory when taking its pointer, and again once the
operation has been submitted, also when that fails. Between the two, an
epoch closed by another task could appear to cover the operation, so only
the task submitting work to a stream closes epochs that count. Epochs
closed by another task, or on a stream that several live tasks submit
work to, therefore don't cover anything. A stream whose task has finished
can be taken over by another one, as happens when streams are recycled.

A pointer taken outside of such an operation, e.g. with `pointer(a)` or
when passing an array to a library call, is only stamped once and can be
used by work that is submitted after its epoch was closed. Such memory
isn't considered complete until it moves to another stream, which waits
for all work submitted to the old one so far.

Recording an event on a stream that is being captured only adds a node
to the graph, so it doesn't close an epoch, and captured operations
don't stamp the memory they use (launching the graph does). Memory used
by captures that CUDA.jl doesn't know about, and memory used on the
special streams, are never considered complete.
When a stream waits for an event, everything it executes afterwards is
ordered after the work the event covers. Memory last used in that work can
then move to that stream without waiting for the stream that used it.
That's what KernelAbstractions' `@spawn` relies on: the spawned task waits
for an event recorded by its parent, but its first use of the parent's
arrays currently also waits for everything the parent queued after the
event.

Have each stream remember, per stream it waited for, up to which epoch its
work is ordered, and don't wait again for memory whose last use is covered
by that. This applies to the memory that could be handed off on the
device anyway, and not to memory whose pointer escaped. The memory remains
marked as waiting on another stream, so that a pointer conversion for a
consumer outside of CUDA.jl, like MPI, synchronizes the new stream. That
only waits for the work the event covered, and not for everything the
other stream queued since, as synchronizing that stream would.
@maleadt
maleadt force-pushed the tb/event-ownership branch from 636c841 to 5903e78 Compare October 7, 2026 04:54

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant