Skip to content

Make CUDA graphs usable - #3326

Open
maleadt wants to merge 12 commits into
tb/stream-poolfrom
tb/graphs
Open

maleadt wants to merge 12 commits into
tb/stream-poolfrom
tb/graphs

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

CUDA graphs let you record a sequence of GPU operations once and then launch the whole sequence with a single, cheap API call. For code made of many short kernels, which is common in Julia because every broadcast is a kernel, the CPU cost of launching each operation often exceeds the time the GPU spends executing it, so graphs can make such code several times faster. CUDA.jl has had capture, instantiate and @captured for a while, but they didn't work with CUDA.jl's GC-based memory management:

This PR reworks the integration so that graphs are safe to use with ordinary Julia code, and adds an API for both simple capture-and-replay and for constructing graphs by hand. It builds on the deferred release of resources from #3318, which is what makes capture compatible with the GC. The commits are self-contained steps, and are best reviewed one at a time.

using CUDA, LinearAlgebra

# a small MLP, evaluated in place
function forward!(hs, Ws, bs)
    for i in eachindex(Ws)
        mul!(hs[i+1], Ws[i], hs[i])
        hs[i+1] .= max.(hs[i+1] .+ bs[i], 0f0)
    end
end

forward!(hs, Ws, bs)                        # warm up: compile kernels, create library handles
exec = instantiate(capture(() -> forward!(hs, Ws, bs)))

for batch in batches
    copyto!(hs[1], batch)                   # inputs: overwrite the captured arrays
    exec()                                  # replay all 8 operations at once
    # ... use hs[end]
end

Allocating code can be captured as well, e.g., y .= W * x .+ b, where W * x allocates a temporary: memory that is allocated during capture is allocated once, outside of the graph, kept alive by it, and overwritten by every launch.

For building graphs by hand, capture! adds captured operations to an existing graph with explicit dependencies, which makes it possible to express concurrency without juggling streams and events. update! changes the parameters of a single operation in an executable graph, e.g., to launch a kernel on different arrays without capturing again:

graph = CuGraph()
a = capture!(graph) do; x .= sin.(x); end
b = capture!(graph) do; y .= cos.(y); end
node = only(capture!(graph; after=[a; b]) do
    @cuda threads=256 combine!(z, x, y)
end)
exec = instantiate(graph)
exec()

update!(exec, node) do                      # retarget the last kernel
    @cuda threads=256 combine!(z2, x, y)
end
exec()

RTX 5080, Julia 1.13. "Submit" is the CPU time to enqueue one iteration, "per iteration" is the throughput when submitting many iterations, and "latency" includes waiting for the result:

workload submit per iteration latency
MLP, 4 layers of 256, batch 16 (4 mul! + 4 broadcasts) eager 49 µs 49 µs 55 µs
graph 8 µs 16 µs 20 µs
RK4, 10 steps of 4096 elements (80 broadcasts, allocating) eager 260 µs 260 µs 265–430 µs
graph 3.5 µs 70 µs 75 µs

Eager execution is CPU-bound: the GPU sits idle between kernels, and the allocating RK4 code also suffers from GC pauses, hence its variable latency. With graphs, the GPU becomes the bottleneck: the CPU only spends half of each MLP iteration submitting work, and 5% of each RK4 iteration. The benchmark suite gains a cuda/graph group tracking eager execution, graph launch and capture of 10 small kernels.

Graphs lease the memory they use. Managed memory gets a lease count, and every way of releasing it (the pool, wrapped host memory, unsafe_free!) goes through a single release gate that postpones the release until the last lease ends. While capturing, take_ownership! records the memory that captured operations use instead of synchronizing or changing its owning stream; the resulting CuGraph, and every CuGraphExec instantiated from it, each hold a lease. Launching an executable graph takes ownership of that memory like a kernel launch does, so other tasks synchronize correctly, and the eventual release is ordered after the last launch.

Capturing no longer disables the GC. Finalizers only publish release requests (#3318). While a capture is in progress, retired resources aren't disposed of (doing so could add frees to the graph, or make prohibited API calls), and the capture drains them when it ends. Captures can't start while a disposal is in progress.

Captures don't get in the way of other tasks. Operations are captured on a dedicated non-blocking stream, which is the capturing task's stream for the duration of the capture, so other tasks can keep using the work the task submitted to its own stream earlier (e.g., an array it produced). capture(f; stream) captures on a specific stream instead. Captures use the relaxed capture mode by default: the driver checks for unsafe API calls per thread rather than per task, so in the stricter modes an unrelated task that ran on the capturing thread while the capturing task was waiting (e.g., on a lock) had its allocations or synchronizations rejected, which invalidated the capture. CUDA.jl checks the operations it performs itself; the stricter modes remain available to debug libraries that don't support capture. Loading a kernel no longer synchronizes the device while a graph is being captured, and a capture that starts while loading a kernel synchronizes the device fails instead of waiting for it.

Allocations during capture happen outside the graph. Graph-owned allocations are only valid between launches, which doesn't fit arrays whose lifetime the GC manages. Instead, memory is allocated from the regular pool on a separate stream, which owns that memory, so whoever uses it next (launching the graph, or another task) waits for the allocation like it would for any other operation. Scalars passed to libraries as CuRefs are initialized right away rather than captured, so cuBLAS calls don't add memcpy nodes, which are expensive to launch alongside kernels.

Host memory is snapshotted. A captured copy from CPU memory would read that memory on every launch, long after it may have been freed. Such copies are staged through pinned memory that the graph keeps alive, giving them snapshot semantics. Copying to CPU memory, waiting for the GPU, or using arrays from a GPUArrays.@cached scope throws a descriptive CaptureError instead of invalidating the capture. Arrays backed by HostMemory can be used to exchange data with the CPU when launching a graph.

Library allocations are handled. Libraries may allocate graph-owned memory during capture, which the driver doesn't free when the executable graph is updated or destroyed. Executable graphs track those allocations, free them when needed (like other memory, deferred while a capture is in progress), and are instantiated with AUTO_FREE_ON_LAUNCH so they can be launched repeatedly. reclaim() also trims the graph memory pool.

Limitations: random numbers generated within kernels (rand()) or by the native CUDA.RNG are the same on every launch, because their seeds are kernel arguments; rand!/CUDA.rand, which use cuRAND, are not affected. Giving captured RNG kernels a device-side seed that a tiny kernel advances at the start of every replay works, but grows KernelState (changing every kernel's ABI) and doesn't cover the GPUArrays RNG, so that's better done separately. Seeding a cuRAND generator synchronizes the device (#3336), so a task that uses cuRAND for the first time, or re-seeds it, while another task is capturing invalidates that capture. Other tasks can't wait for work submitted to an explicitly passed capture stream before the capture until it ends. Tasks spawned during capture run their operations eagerly, outside the graph. Graphs captured directly through the driver API aren't integrated with the memory management.

Tested on an RTX 5080 with Julia 1.12 (core/graph, core/resource_graph).

Fixes #3310, fixes #1230, fixes #2691. #1229 now results in a descriptive error.

@maleadt
maleadt added this pull request to stack #3319 October 3, 2026 13:05
@vchuravy

vchuravy commented Oct 3, 2026

Copy link
Copy Markdown
Member

What happens under task parallelism?

@maleadt
maleadt removed this pull request from stack #3319 October 3, 2026 15:49
@maleadt
maleadt changed the base branch from tb/deferred-release to tb/stream-pool October 3, 2026 15:49
@maleadt
maleadt added this pull request to stack #3327 October 3, 2026 15:49
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 88.28996% with 63 lines in your changes missing coverage. Please review.
✅ Project coverage is 87.15%. Comparing base (7623dd4) to head (1f0818f).

Files with missing lines Patch % Lines
CUDACore/src/graph.jl 86.51% 60 Missing ⚠️
CUDACore/src/resources/managed.jl 90.47% 2 Missing ⚠️
CUDACore/lib/cudadrv/graph.jl 88.88% 1 Missing ⚠️
Additional details and impacted files
@@                Coverage Diff                 @@
##           tb/stream-pool    #3326      +/-   ##
==================================================
+ Coverage           87.09%   87.15%   +0.06%     
==================================================
  Files                 194      195       +1     
  Lines               19310    19749     +439     
==================================================
+ Hits                16818    17213     +395     
- Misses               2492     2536      +44     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt marked this pull request as ready for review October 5, 2026 04:51
@maleadt
maleadt removed this pull request from stack #3327 October 5, 2026 08:33
@maleadt
maleadt added this pull request to stack #3338 October 5, 2026 08:33
@maleadt

maleadt commented Oct 5, 2026 •

Copy link
Copy Markdown
Member Author

🤖 Good question; it wasn't handled well. Concurrent captures, launching a graph from several tasks, and unrelated tasks using the GPU on other threads during a capture all worked, but:

  • thread-local capture mode protects an OS thread, not a task, so tasks that ran on the capturing thread while the capture was blocked (e.g., on a lock) failed to allocate or synchronize, and invalidated the capture;
  • another task loading a kernel, or using an array the capturing task had produced, got a CaptureError;
  • tasks spawned inside capture executed their operations eagerly instead of capturing them.

The first two are fixed here: captures now use the relaxed mode and a dedicated stream, and kernel loading no longer synchronizes the device while capturing. Capturing the operations of spawned tasks, e.g. from libraries that parallelize internally, is #3337: they inherit the capture through a scoped value and submit to the same stream. What remains is a cuRAND issue (#3336): seeding a generator synchronizes the device, which breaks captures elsewhere.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 1f0818f Previous: e924726 Ratio
array/accumulate/Float32/1d 99207 ns 99107 ns 1.00
array/accumulate/Float32/dims=1 73377 ns 72803 ns 1.01
array/accumulate/Float32/dims=1L 1587880 ns 1589176 ns 1.00
array/accumulate/Float32/dims=2 138529 ns 138542 ns 1.00
array/accumulate/Float32/dims=2L 656451 ns 655298 ns 1.00
array/accumulate/Int64/1d 117845 ns 117327 ns 1.00
array/accumulate/Int64/dims=1 77163 ns 77135 ns 1.00
array/accumulate/Int64/dims=1L 1702448 ns 1698128 ns 1.00
array/accumulate/Int64/dims=2 153133 ns 150131 ns 1.02
array/accumulate/Int64/dims=2L 987317 ns 986795 ns 1.00
array/broadcast 16193 ns 16029 ns 1.01
array/broadcast launch 7493 ns 7358.25 ns 1.02
array/construct 939.6666666666666 ns 959.9 ns 0.98
array/copy 17111 ns 16576 ns 1.03
array/copyto!/cpu_to_gpu 206523 ns 207398 ns 1.00
array/copyto!/gpu_to_cpu 240228 ns 239800 ns 1.00
array/copyto!/gpu_to_gpu 10104.666666666666 ns 8720.666666666666 ns 1.16
array/iteration/findall/bool 133483 ns 129599 ns 1.03
array/iteration/findall/int 145165 ns 140825 ns 1.03
array/iteration/findfirst/bool 71106 ns 67856 ns 1.05
array/iteration/findfirst/int 73664 ns 69203 ns 1.06
array/iteration/findmin/1d 65811 ns 59849 ns 1.10
array/iteration/findmin/2d 97397 ns 97040 ns 1.00
array/iteration/logical 190846 ns 181916 ns 1.05
array/iteration/scalar 55692 ns 58365 ns 0.95
array/permutedims/2d 46387 ns 45704 ns 1.01
array/permutedims/3d 47729 ns 47728 ns 1.00
array/permutedims/4d 48685 ns 48656 ns 1.00
array/random/rand/Float32 11075 ns 10851 ns 1.02
array/random/rand/Int64 18263 ns 18300 ns 1.00
array/random/rand!/Float32 7805.75 ns 7842.666666666667 ns 1.00
array/random/rand!/Int64 16886 ns 16741 ns 1.01
array/random/randn/Float32 32236 ns 32264 ns 1.00
array/random/randn!/Float32 23339 ns 24124 ns 0.97
array/reductions/mapreduce/Float32/1d 34671 ns 32799 ns 1.06
array/reductions/mapreduce/Float32/dims=1 37806 ns 37628 ns 1.00
array/reductions/mapreduce/Float32/dims=1L 51582 ns 51393 ns 1.00
array/reductions/mapreduce/Float32/dims=2 55539 ns 55676 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 67873 ns 67890 ns 1.00
array/reductions/mapreduce/Int64/1d 41154 ns 39071 ns 1.05
array/reductions/mapreduce/Int64/dims=1 40962 ns 40979 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 89227 ns 89303 ns 1.00
array/reductions/mapreduce/Int64/dims=2 56953 ns 57656 ns 0.99
array/reductions/mapreduce/Int64/dims=2L 83905 ns 84117 ns 1.00
array/reductions/reduce/Float32/1d 35071 ns 32726 ns 1.07
array/reductions/reduce/Float32/dims=1 38125 ns 37671 ns 1.01
array/reductions/reduce/Float32/dims=1L 51378 ns 51035 ns 1.01
array/reductions/reduce/Float32/dims=2 55520 ns 55445 ns 1.00
array/reductions/reduce/Float32/dims=2L 67672 ns 68151 ns 0.99
array/reductions/reduce/Int64/1d 41299 ns 39207 ns 1.05
array/reductions/reduce/Int64/dims=1 40982 ns 40842 ns 1.00
array/reductions/reduce/Int64/dims=1L 88814 ns 88981 ns 1.00
array/reductions/reduce/Int64/dims=2 57004 ns 58095 ns 0.98
array/reductions/reduce/Int64/dims=2L 84101 ns 84484 ns 1.00
array/reverse/1d 17351 ns 17482 ns 0.99
array/reverse/1dL 70147 ns 70086 ns 1.00
array/reverse/1dL_inplace 67794 ns 67824 ns 1.00
array/reverse/1d_inplace 8999 ns 9051.333333333334 ns 0.99
array/reverse/2d 20137 ns 20350 ns 0.99
array/reverse/2dL 73992 ns 73763 ns 1.00
array/reverse/2dL_inplace 67445 ns 67393 ns 1.00
array/reverse/2d_inplace 10179 ns 10090 ns 1.01
array/sorting/1d 2638416 ns 2646314 ns 1.00
array/sorting/2d 1018095 ns 1018011 ns 1.00
array/sorting/by 3158308 ns 3158564 ns 1.00
cuda/graph/capture 27814 ns
cuda/graph/eager 38629 ns
cuda/graph/launch 14172 ns
cuda/synchronization/context/auto 7900.333333333333 ns 6775.2 ns 1.17
cuda/synchronization/context/blocking 888.3409090909091 ns 808.9775280898876 ns 1.10
cuda/synchronization/context/nonblocking 7930 ns 6745.2 ns 1.18
cuda/synchronization/stream/auto 719.6349206349206 ns 707.3802816901408 ns 1.02
cuda/synchronization/stream/blocking 942.0434782608696 ns 862.0714285714286 ns 1.09
cuda/synchronization/stream/nonblocking 8386.333333333334 ns 7079.5 ns 1.18
integration/byval/reference 148246 ns 148360 ns 1.00
integration/byval/slices=1 149529 ns 149203 ns 1.00
integration/byval/slices=2 292421 ns 292017 ns 1.00
integration/byval/slices=3 435797 ns 434881 ns 1.00
integration/cudadevrt 105544 ns 105387 ns 1.00
integration/volumerhs 9145951 ns 9145562 ns 1.00
kernel/indexing 13320 ns 13190 ns 1.01
kernel/indexing_checked 13800 ns 13943 ns 0.99
kernel/launch 2430.1111111111113 ns 2472 ns 0.98
kernel/occupancy 953.4090909090909 ns 938.1739130434783 ns 1.02
kernel/rand 15048 ns 14010 ns 1.07
latency/import 4313343915 ns 4302362850 ns 1.00
latency/precompile 5103507016 ns 5082487750 ns 1.00
latency/ttfp 4827505049 ns 4807353177 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Memory can be in use beyond the lifetime of the object that owns it,
e.g., by a graph that captured operations on an array. Add a lease count
to `Managed` memory, and route every way of releasing managed memory
(the pool, and wrapped host memory) through `release`, which postpones
the release until the last lease has ended.
Capturing used to disable the GC for the whole process, as freeing
memory while capturing breaks the capture. With deferred release,
finalizers don't call into CUDA anymore, so instead of disabling the GC,
don't dispose of retired resources while a capture is in progress, and
drain them when it ends. Captures start under the drain lock, so they
can't race with disposal that's already in progress.

The graph functionality moves out of the driver wrappers, and captures
now:
- use the relaxed capture mode by default. The driver's checks for
  unsafe API calls apply per thread, not per task, so in the global or
  thread-local mode they also rejected operations of unrelated tasks
  that ran on the capturing thread while the capturing task was waiting,
  invalidating the capture. CUDA.jl checks the operations it performs
  itself, and the stricter modes remain available;
- keep the capturing task on its thread in those stricter modes, as such
  captures need to end on the thread they started on;
- end the capture on every exit path, and report the original error;
- report waiting for the GPU while capturing with a descriptive
  `CaptureError`.
maleadt added 10 commits October 6, 2026 15:37
Executable graphs are now callable, and can be uploaded to the device
ahead of their first launch. Graphs show their contents, can be rendered
with Graphviz, and their nodes can be queried.
Graphs used to keep none of the memory their operations use alive, so
launching a graph after one of its arrays had been freed accessed freed
memory. Now, operations that use memory while capturing record that
memory instead of synchronizing with its owner, and the resulting graph,
as well as every executable graph instantiated from it, holds a lease on
it. Launching an executable graph takes ownership of that memory like
launching a kernel does, so that other tasks synchronize with it, and so
that releasing the memory is ordered after the last launch.

Arrays allocated from a `GPUArrays.@cached` scope are rejected while
capturing, as the cache reuses their memory regardless of what still
uses it.
Memory allocated in stream order while capturing becomes owned by the
graph, valid only between launches, which doesn't fit arrays whose
lifetime is managed by the GC (#1230). Instead, allocate such memory
from the pool on a separate stream, and keep it alive using leases. That
stream owns the memory, so whoever uses it next (launching the graph, or
another task) waits for the allocation through the usual ownership
mechanism. Every launch of the graph then overwrites the arrays that
were allocated during capture.

Scalar references (`CuRef`) are initialized right away instead of being
captured, as memory copies are relatively expensive to launch as part of
a graph. This also fixes capturing library calls with scalar arguments
(#2691).
A captured copy from host memory reads that memory every time the graph
is launched, long after the memory may have been freed. Copy such data
into pinned memory that the graph keeps alive, giving captured copies
snapshot semantics. Copying to host memory can't be made safe the same
way, so it results in a `CaptureError`; arrays backed by `HostMemory`
can be used instead.
Libraries may allocate memory in stream order while being captured,
which becomes owned by the graph. Launching such a graph again requires
freeing that memory first, so instantiate graphs with
`AUTO_FREE_ON_LAUNCH`, and free the outstanding allocations when
updating or destroying an executable graph, as the driver doesn't.
Reclaiming memory now also trims the memory that graphs keep cached.
`capture!` adds captured operations to an existing graph, depending on
given nodes, which makes it possible to construct graphs with operations
that execute concurrently without juggling streams and events. `update!`
changes the parameters of a single operation in an executable graph,
e.g., to launch a kernel with different arguments, without capturing the
graph again.
Capturing on the task's stream made it impossible for other tasks to
wait for work the capturing task had submitted before the capture, e.g.,
to use an array it produced. Instead, check out a non-blocking stream
that's only used for capturing, and make it the task's stream for the
duration of the capture.
capture and capture! now accept a stream keyword argument, to capture on
that stream instead of on a dedicated one. It needs to be a non-blocking
stream of the current context, which is reserved for the capture, and is
the task's stream while capturing.
The workaround for module loading under memory pressure synchronized the
context unless the current task was capturing, which invalidated
captures by other tasks. Skip it whenever a graph is being captured.
Captures that start during the synchronization fail instead of waiting
for it, as they do while reclaiming memory, since the awaited work may
depend on the capturing task.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Freeing memory from finalizers during graph capture @captured with scalar arguments fails with CUDA 5.7 Graph memory management

2 participants