Repository navigation
Keep wrapped CPU arrays alive in unsafe_wrap(CuArray, ...) - #3308
Merged
Merged
Conversation
`unsafe_wrap(CuArray, a::Array)` did not retain `a`, so the CuArray could outlive the memory it wraps once the Array was garbage collected, both for HMM-backed unified memory and for registered host memory (whose finalizer would then unregister freed memory). Capture the Array in the DataRef finalizer so that it lives as long as the wrapper and any views derived from it. Wrapping a raw pointer keeps its existing semantics. Releasing the wrapper now also waits for outstanding work on it before letting go of the Array or unregistering the memory, so dropping the last reference while a kernel is still using it is safe. If that wait fails, the memory is leaked rather than released while it may still be in use. Also document that the reverse, `unsafe_wrap(Array, ::CuArray)`, does not keep the CuArray alive, and that captured graphs don't keep wrapped memory alive.
Unregistering host memory already waits for outstanding work using it, so do that directly, with the thread's capture mode relaxed so that it neither fails nor invalidates a graph capture that is in progress. Unified memory has no such operation, so have the stream signal when the work is done, by launching a host function. The task waiting for that signal keeps the Array alive, so failing to signal it leaks the memory without needing a global list. When the stream is being captured, retry after the capture has finished.
CUDA.capture disables the GC, so re-enable it while collecting.
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
A `let` at the top level of a test loop doesn't make the wrapper unreachable, so it was never collected and the test only passed if the kernel happened to finish before the GC did.
Both memory kinds now share one mechanism: the wrapper's finalizer wakes a task, created along with the wrapper (so it isn't affected by cancellation), which waits for outstanding work on the stream that last used the memory, unregisters page-locked memory, and releases the Array. An explicit unsafe_free! waits for that task, so the Array can be wrapped again right away. This no longer relies on cuMemHostUnregister implicitly waiting for the GPU, launches nothing on the stream, doesn't block in finalizers, and removes the capture-mode swap and the polling task. GC during a capture is left unsupported, like elsewhere in CUDA.jl, so the test that re-enabled it (and freed unrelated memory on the capturing stream) is removed. The ordering tests now use explicit frees, reading the kernel's result from wrapped host memory, which doesn't depend on GC timing.
maleadt
enabled auto-merge
September 30, 2026 10:48
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #3308 +/- ##
==========================================
+ Coverage 85.93% 86.06% +0.12%
==========================================
Files 187 187
Lines 19025 19063 +38
==========================================
+ Hits 16349 16406 +57
+ Misses 2676 2657 -19 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
unsafe_wrap(CuArray, a::Array)gives the GPU direct access to a CPU array, without copying: with HMM the GPU reads the pageable memory directly, and otherwise the memory is page-locked. Until now, the resultingCuArraydid not keepaalive, so the obvious way to use it was a use-after-free waiting to happen:With this PR, a
CuArraycreated from anArraykeeps it alive for as long as the wrapper, or anything derived from it, exists, and until the GPU is done with it. That's also what the OpenCL.jl, Metal.jl and AMDGPU.jl versions of this functionality do (JuliaGPU/OpenCL.jl#513, JuliaGPU/Metal.jl#991, JuliaGPU/AMDGPU.jl#1116). Wrapping a raw pointer is unchanged: keeping that memory alive is still up to you.Releasing the wrapper waits for outstanding work on the stream that last used the memory, and only then unregisters page-locked memory and drops the reference to the array. The same happens for HMM and page-locked memory, so there's no reliance on
cuMemHostUnregisterimplicitly waiting for the GPU (which isn't documented). Waiting requires switching tasks, which finalizers can't do, so each wrapper comes with a task (anAsyncConditioncallback, which is shielded from cancellation) that the finalizer wakes up. When the wrapper is freed explicitly withunsafe_free!, that call waits for the task, so the array can be wrapped again right away. If the stream is being captured when the memory is released, the task waits for the capture to end first. Nothing is launched on the stream and nothing blocks in a finalizer.The docstring now also has a prominent warning that the other direction,
unsafe_wrap(Array, ::CuArray), does not keep theCuArrayalive. It also notes that CUDA graphs capturing operations on a wrapped array don't keep that array alive either.Tested on an RTX 5080, which supports HMM, so both the HMM and the page-locked paths were exercised. The tests check that the kernel using a wrapped array has finished by the time the array is released, including when the stream it ran on was destroyed in the meantime. Versions of this PR that don't wait fail them.