Repository navigation
Fix waiting for the GPU before releasing wrapped memory - #3320
Merged
Merged
Conversation
The release of wrapped system memory still called `nonblocking_synchronize`, which #3315 removed. Waiting failed with an UndefVarError that was only logged, after which the memory was unregistered and the Array released while the GPU could still be using it. Use `cooperative_wait` instead. Also wait and unregister with a relaxed capture mode: freeing a wrapper while another stream is being captured, e.g. by the task doing the freeing, otherwise invalidated that capture.
maleadt
added this pull request to stack #3321
October 2, 2026 13:34
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
Releasing wrapped memory last used on a special or destroyed stream waits for the entire context, which is prohibited while any of its streams is being captured, regardless of the capture mode. Postpone that until captures made with `capture` have ended.
maleadt
force-pushed
the
tb/wrap-release-capture
branch
from
October 3, 2026 07:58
c1bd27b to
f1faae0
Compare
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #3320 +/- ##
===========================================
+ Coverage 39.40% 86.23% +46.83%
===========================================
Files 187 187
Lines 18966 18997 +31
===========================================
+ Hits 7474 16383 +8909
+ Misses 11492 2614 -8878 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Wrapping a CPU array with
unsafe_wrap(CuArray, a)keepsaalive, and page-locked if needed, until the GPU is done with it (#3308). After the wrapper is freed, a task waits for the GPU work that used the memory, and only then unregisters it and lets go ofa. This PR fixes two ways in which that release goes wrong.The release no longer waits for the GPU on main
#3315 replaced CUDA.jl's nonblocking synchronization with GPUToolbox's
cooperative_wait, and removednonblocking_synchronize. The release code from #3308 still called it, and the two PRs were merged around the same time. So on main, releasing a wrapper that the GPU is still using logsand then releases the memory anyway, while the kernel may still be reading it. The existing tests catch this, but only when the kernel is still running at release time: two
@test outh[] == sum(1:1024)checks intest/core/array.jlfail on main here. The wait now usescooperative_wait. Because a failed release is only logged, the tests also check that releasing the memory logs no errors.Freeing a wrapper during a graph capture invalidated it
Consider a task that captures a graph and, inside the capture, frees a wrapper that another task last used:
To release
b, CUDA.jl has to wait for that other stream and then unregister the memory. Even though neither involves the stream being captured, CUDA forbids both calls while a capture in the default (global) mode is in progress, so the capture got invalidated:These calls now use the relaxed capture mode, which allows them because they don't touch the stream being captured. That isn't enough if the wrapper was last used on one of the special streams (default, legacy or per-thread), or on a stream that has been destroyed since. Releasing it then means waiting for the whole context, which isn't allowed during a capture in any mode. In that case the release is postponed until captures made with
CUDA.capturehave finished. The explicitunsafe_free!doesn't wait for it in that case. Captures started directly through the driver API can't be detected this way.Found while porting the review fixes of JuliaGPU/AMDGPU.jl#1116, which has the same release logic. #3317 is stacked on top of this PR, since both change
captureand the release logic.Tested on an RTX 5080:
core/arrayandcore/cudadrvpass. The new tests fail without these changes, including the destroyed-stream case when the postponement is disabled.