Skip to content

Keep wrapped CPU arrays alive in unsafe_wrap(CuArray, ...) - #3308

Merged
maleadt merged 5 commits into
mainfrom
tb/unsafe_wrap_lifetime
Sep 30, 2026
Merged

maleadt merged 5 commits into
mainfrom
tb/unsafe_wrap_lifetime

Conversation

@maleadt

@maleadt maleadt commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

unsafe_wrap(CuArray, a::Array) gives the GPU direct access to a CPU array, without copying: with HMM the GPU reads the pageable memory directly, and otherwise the memory is page-locked. Until now, the resulting CuArray did not keep a alive, so the obvious way to use it was a use-after-free waiting to happen:

b = unsafe_wrap(CuArray, rand(Float32, 1024))   # nothing references the Array anymore
GC.gc()                                         # ...so it can be collected here
b .+= 1                                         # and the GPU writes to freed memory

With this PR, a CuArray created from an Array keeps it alive for as long as the wrapper, or anything derived from it, exists, and until the GPU is done with it. That's also what the OpenCL.jl, Metal.jl and AMDGPU.jl versions of this functionality do (JuliaGPU/OpenCL.jl#513, JuliaGPU/Metal.jl#991, JuliaGPU/AMDGPU.jl#1116). Wrapping a raw pointer is unchanged: keeping that memory alive is still up to you.

Releasing the wrapper waits for outstanding work on the stream that last used the memory, and only then unregisters page-locked memory and drops the reference to the array. The same happens for HMM and page-locked memory, so there's no reliance on cuMemHostUnregister implicitly waiting for the GPU (which isn't documented). Waiting requires switching tasks, which finalizers can't do, so each wrapper comes with a task (an AsyncCondition callback, which is shielded from cancellation) that the finalizer wakes up. When the wrapper is freed explicitly with unsafe_free!, that call waits for the task, so the array can be wrapped again right away. If the stream is being captured when the memory is released, the task waits for the capture to end first. Nothing is launched on the stream and nothing blocks in a finalizer.

The docstring now also has a prominent warning that the other direction, unsafe_wrap(Array, ::CuArray), does not keep the CuArray alive. It also notes that CUDA graphs capturing operations on a wrapped array don't keep that array alive either.

Tested on an RTX 5080, which supports HMM, so both the HMM and the page-locked paths were exercised. The tests check that the kernel using a wrapped array has finished by the time the array is released, including when the stream it ran on was destroyed in the meantime. Versions of this PR that don't wait fail them.

`unsafe_wrap(CuArray, a::Array)` did not retain `a`, so the CuArray could outlive the
memory it wraps once the Array was garbage collected, both for HMM-backed unified memory
and for registered host memory (whose finalizer would then unregister freed memory).
Capture the Array in the DataRef finalizer so that it lives as long as the wrapper and
any views derived from it. Wrapping a raw pointer keeps its existing semantics.

Releasing the wrapper now also waits for outstanding work on it before letting go of the
Array or unregistering the memory, so dropping the last reference while a kernel is still
using it is safe. If that wait fails, the memory is leaked rather than released while it
may still be in use.

Also document that the reverse, `unsafe_wrap(Array, ::CuArray)`, does not keep the
CuArray alive, and that captured graphs don't keep wrapped memory alive.
Unregistering host memory already waits for outstanding work using it, so
do that directly, with the thread's capture mode relaxed so that it
neither fails nor invalidates a graph capture that is in progress.

Unified memory has no such operation, so have the stream signal when the
work is done, by launching a host function. The task waiting for that
signal keeps the Array alive, so failing to signal it leaks the memory
without needing a global list. When the stream is being captured, retry
after the capture has finished.
CUDA.capture disables the GC, so re-enable it while collecting.
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 27edd84 Previous: c41ddb2 Ratio
array/accumulate/Float32/1d 99525 ns 99024 ns 1.01
array/accumulate/Float32/dims=1 71520 ns 72145 ns 0.99
array/accumulate/Float32/dims=1L 1588281 ns 1588461 ns 1.00
array/accumulate/Float32/dims=2 137148 ns 137848 ns 0.99
array/accumulate/Float32/dims=2L 655689 ns 654636 ns 1.00
array/accumulate/Int64/1d 119057 ns 117263 ns 1.02
array/accumulate/Int64/dims=1 75387 ns 76091 ns 0.99
array/accumulate/Int64/dims=1L 1706235 ns 1697362 ns 1.01
array/accumulate/Int64/dims=2 149291 ns 149255 ns 1.00
array/accumulate/Int64/dims=2L 987046 ns 986192 ns 1.00
array/broadcast 16660 ns 15428 ns 1.08
array/broadcast launch 8168.333333333333 ns 6869 ns 1.19
array/construct 878.6862745098039 ns 870.7142857142857 ns 1.01
array/copy 16518 ns 16625 ns 0.99
array/copyto!/cpu_to_gpu 206802 ns 206218 ns 1.00
array/copyto!/gpu_to_cpu 239760 ns 239720 ns 1.00
array/copyto!/gpu_to_gpu 9958.333333333334 ns 10239 ns 0.97
array/iteration/findall/bool 133745 ns 131631 ns 1.02
array/iteration/findall/int 145056 ns 143514 ns 1.01
array/iteration/findfirst/bool 68550 ns 68312 ns 1.00
array/iteration/findfirst/int 70518 ns 70287 ns 1.00
array/iteration/findmin/1d 64340 ns 60167 ns 1.07
array/iteration/findmin/2d 97864 ns 96150 ns 1.02
array/iteration/logical 186795 ns 178893 ns 1.04
array/iteration/scalar 60229 ns 63853 ns 0.94
array/permutedims/2d 47465 ns 45346 ns 1.05
array/permutedims/3d 48821 ns 46982 ns 1.04
array/permutedims/4d 50044 ns 47789 ns 1.05
array/random/rand/Float32 11391 ns 11572 ns 0.98
array/random/rand/Int64 19835 ns 18453 ns 1.07
array/random/rand!/Float32 7625.75 ns 7761.5 ns 0.98
array/random/rand!/Int64 17531 ns 16118 ns 1.09
array/random/randn/Float32 32646 ns 32223 ns 1.01
array/random/randn!/Float32 25987 ns 23621 ns 1.10
array/reductions/mapreduce/Float32/1d 31485 ns 32045 ns 0.98
array/reductions/mapreduce/Float32/dims=1 37062 ns 37283 ns 0.99
array/reductions/mapreduce/Float32/dims=1L 50735 ns 50954 ns 1.00
array/reductions/mapreduce/Float32/dims=2 54848 ns 54960 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 66512 ns 66739 ns 1.00
array/reductions/mapreduce/Int64/1d 38821 ns 39254 ns 0.99
array/reductions/mapreduce/Int64/dims=1 39837 ns 39772 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 88307 ns 88484 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57212 ns 57117 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 82978 ns 82793 ns 1.00
array/reductions/reduce/Float32/1d 31763 ns 32298 ns 0.98
array/reductions/reduce/Float32/dims=1 37217 ns 36873 ns 1.01
array/reductions/reduce/Float32/dims=1L 50370 ns 50474 ns 1.00
array/reductions/reduce/Float32/dims=2 54843 ns 54877 ns 1.00
array/reductions/reduce/Float32/dims=2L 66636 ns 66929 ns 1.00
array/reductions/reduce/Int64/1d 39269 ns 38693 ns 1.01
array/reductions/reduce/Int64/dims=1 39897 ns 39824 ns 1.00
array/reductions/reduce/Int64/dims=1L 88333 ns 88360 ns 1.00
array/reductions/reduce/Int64/dims=2 57223 ns 57104 ns 1.00
array/reductions/reduce/Int64/dims=2L 83005 ns 82794 ns 1.00
array/reverse/1d 16556 ns 16811 ns 0.98
array/reverse/1dL 69217 ns 69626 ns 0.99
array/reverse/1dL_inplace 67214 ns 67260 ns 1.00
array/reverse/1d_inplace 8713.333333333334 ns 8336 ns 1.05
array/reverse/2d 19590 ns 19529 ns 1.00
array/reverse/2dL 72990 ns 73101 ns 1.00
array/reverse/2dL_inplace 66999 ns 67223 ns 1.00
array/reverse/2d_inplace 10100 ns 12107 ns 0.83
array/sorting/1d 2655627 ns 2646351 ns 1.00
array/sorting/2d 1013605 ns 1018059 ns 1.00
array/sorting/by 3170919 ns 3173715 ns 1.00
cuda/synchronization/context/auto 1018.8 ns 1015.9 ns 1.00
cuda/synchronization/context/blocking 794.7525773195877 ns 787.757281553398 ns 1.01
cuda/synchronization/context/nonblocking 5713.5 ns 5696.333333333333 ns 1.00
cuda/synchronization/stream/auto 854.5166666666667 ns 866.1666666666666 ns 0.99
cuda/synchronization/stream/blocking 671.1948051948052 ns 661.1446540880503 ns 1.02
cuda/synchronization/stream/nonblocking 5651.333333333333 ns 5718.666666666667 ns 0.99
integration/byval/reference 147753 ns 147659 ns 1.00
integration/byval/slices=1 148979 ns 148771 ns 1.00
integration/byval/slices=2 291542 ns 291271 ns 1.00
integration/byval/slices=3 434588 ns 434742 ns 1.00
integration/cudadevrt 104791 ns 104894 ns 1.00
integration/volumerhs 9143690 ns 9147760 ns 1.00
kernel/indexing 12850 ns 13026 ns 0.99
kernel/indexing_checked 13422 ns 13436 ns 1.00
kernel/launch 2088.4444444444443 ns 2109.3333333333335 ns 0.99
kernel/occupancy 669.496855345912 ns 699.2582781456954 ns 0.96
kernel/rand 16929 ns 13938 ns 1.21
latency/import 4249196255 ns 4239937392 ns 1.00
latency/precompile 5024722530 ns 5025064161 ns 1.00
latency/ttfp 4705138566 ns 4726157535 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

A `let` at the top level of a test loop doesn't make the wrapper
unreachable, so it was never collected and the test only passed if the
kernel happened to finish before the GC did.
Both memory kinds now share one mechanism: the wrapper's finalizer wakes
a task, created along with the wrapper (so it isn't affected by
cancellation), which waits for outstanding work on the stream that last
used the memory, unregisters page-locked memory, and releases the Array.
An explicit unsafe_free! waits for that task, so the Array can be wrapped
again right away. This no longer relies on cuMemHostUnregister implicitly
waiting for the GPU, launches nothing on the stream, doesn't block in
finalizers, and removes the capture-mode swap and the polling task.

GC during a capture is left unsupported, like elsewhere in CUDA.jl, so
the test that re-enabled it (and freed unrelated memory on the capturing
stream) is removed. The ordering tests now use explicit frees, reading
the kernel's result from wrapped host memory, which doesn't depend on
GC timing.
@maleadt
maleadt enabled auto-merge September 30, 2026 10:48
@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.00000% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 86.06%. Comparing base (997e2e0) to head (27edd84).
⚠️ Report is 6 commits behind head on main.

Files with missing lines Patch % Lines
CUDACore/src/array.jl 92.00% 4 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3308      +/-   ##
==========================================
+ Coverage   85.93%   86.06%   +0.12%     
==========================================
  Files         187      187              
  Lines       19025    19063      +38     
==========================================
+ Hits        16349    16406      +57     
+ Misses       2676     2657      -19     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt disabled auto-merge September 30, 2026 19:31
@maleadt
maleadt merged commit 243d8c8 into main Sep 30, 2026
1 check passed
@maleadt
maleadt deleted the tb/unsafe_wrap_lifetime branch September 30, 2026 19:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant