Skip to content

Fix waiting for the GPU before releasing wrapped memory - #3320

Merged
maleadt merged 2 commits into
mainfrom
tb/wrap-release-capture
Oct 3, 2026
Merged

maleadt merged 2 commits into
mainfrom
tb/wrap-release-capture

Conversation

@maleadt

@maleadt maleadt commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Wrapping a CPU array with unsafe_wrap(CuArray, a) keeps a alive, and page-locked if needed, until the GPU is done with it (#3308). After the wrapper is freed, a task waits for the GPU work that used the memory, and only then unregisters it and lets go of a. This PR fixes two ways in which that release goes wrong.

The release no longer waits for the GPU on main

#3315 replaced CUDA.jl's nonblocking synchronization with GPUToolbox's cooperative_wait, and removed nonblocking_synchronize. The release code from #3308 still called it, and the two PRs were merged around the same time. So on main, releasing a wrapper that the GPU is still using logs

┌ Error: Error while waiting for GPU work on wrapped system memory
│   exception =
│    UndefVarError: `nonblocking_synchronize` not defined in `CUDACore`

and then releases the memory anyway, while the kernel may still be reading it. The existing tests catch this, but only when the kernel is still running at release time: two @test outh[] == sum(1:1024) checks in test/core/array.jl fail on main here. The wait now uses cooperative_wait. Because a failed release is only logged, the tests also check that releasing the memory logs no errors.

Freeing a wrapper during a graph capture invalidated it

Consider a task that captures a graph and, inside the capture, frees a wrapper that another task last used:

graph = CUDA.capture() do
    CUDA.unsafe_free!(b)   # `b` was last used on another task's stream
    c .+= 1
end

To release b, CUDA.jl has to wait for that other stream and then unregister the memory. Even though neither involves the stream being captured, CUDA forbids both calls while a capture in the default (global) mode is in progress, so the capture got invalidated:

Returning 900 (CUDA_ERROR_STREAM_CAPTURE_UNSUPPORTED) from cuStreamSynchronize
Returning 900 (CUDA_ERROR_STREAM_CAPTURE_UNSUPPORTED) from cuMemHostUnregister
capture failed: CUDA error: operation failed due to a previous error during capture (code 901, ERROR_STREAM_CAPTURE_INVALIDATED)

These calls now use the relaxed capture mode, which allows them because they don't touch the stream being captured. That isn't enough if the wrapper was last used on one of the special streams (default, legacy or per-thread), or on a stream that has been destroyed since. Releasing it then means waiting for the whole context, which isn't allowed during a capture in any mode. In that case the release is postponed until captures made with CUDA.capture have finished. The explicit unsafe_free! doesn't wait for it in that case. Captures started directly through the driver API can't be detected this way.

Found while porting the review fixes of JuliaGPU/AMDGPU.jl#1116, which has the same release logic. #3317 is stacked on top of this PR, since both change capture and the release logic.

Tested on an RTX 5080: core/array and core/cudadrv pass. The new tests fail without these changes, including the destroyed-stream case when the postponement is disabled.

The release of wrapped system memory still called `nonblocking_synchronize`,
which #3315 removed. Waiting failed with an UndefVarError that was only logged,
after which the memory was unregistered and the Array released while the GPU
could still be using it. Use `cooperative_wait` instead.

Also wait and unregister with a relaxed capture mode: freeing a wrapper while
another stream is being captured, e.g. by the task doing the freeing, otherwise
invalidated that capture.
@maleadt
maleadt added this pull request to stack #3321 October 2, 2026 13:34
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: f1faae0 Previous: a3202f2 Ratio
array/accumulate/Float32/1d 99002 ns 99028 ns 1.00
array/accumulate/Float32/dims=1 73164 ns 72921 ns 1.00
array/accumulate/Float32/dims=1L 1589686 ns 1588681 ns 1.00
array/accumulate/Float32/dims=2 138112 ns 138806 ns 1.00
array/accumulate/Float32/dims=2L 656780 ns 655618 ns 1.00
array/accumulate/Int64/1d 117696 ns 118327 ns 0.99
array/accumulate/Int64/dims=1 77922 ns 76942 ns 1.01
array/accumulate/Int64/dims=1L 1699890 ns 1697861 ns 1.00
array/accumulate/Int64/dims=2 151639 ns 151829 ns 1.00
array/accumulate/Int64/dims=2L 987820 ns 986767 ns 1.00
array/broadcast 16247 ns 15832 ns 1.03
array/broadcast launch 7374.25 ns 6879.4 ns 1.07
array/construct 922.6666666666666 ns 866.433962264151 ns 1.06
array/copy 16469 ns 16688 ns 0.99
array/copyto!/cpu_to_gpu 208532 ns 208406 ns 1.00
array/copyto!/gpu_to_cpu 241321 ns 241247 ns 1.00
array/copyto!/gpu_to_gpu 10181.666666666666 ns 8802.333333333334 ns 1.16
array/iteration/findall/bool 129739 ns 130199 ns 1.00
array/iteration/findall/int 141963 ns 138662 ns 1.02
array/iteration/findfirst/bool 68895 ns 67562 ns 1.02
array/iteration/findfirst/int 70170 ns 69121 ns 1.02
array/iteration/findmin/1d 61723 ns 59127 ns 1.04
array/iteration/findmin/2d 97285 ns 96956 ns 1.00
array/iteration/logical 184860 ns 180975 ns 1.02
array/iteration/scalar 59617 ns 58859 ns 1.01
array/permutedims/2d 46304 ns 46242 ns 1.00
array/permutedims/3d 48005 ns 46855 ns 1.02
array/permutedims/4d 48702 ns 48368 ns 1.01
array/random/rand/Float32 10542 ns 11809 ns 0.89
array/random/rand/Int64 18322 ns 18752 ns 0.98
array/random/rand!/Float32 7790.5 ns 7829.5 ns 1.00
array/random/rand!/Int64 16891 ns 15615 ns 1.08
array/random/randn/Float32 32485 ns 32695 ns 0.99
array/random/randn!/Float32 23170 ns 23566 ns 0.98
array/reductions/mapreduce/Float32/1d 33237 ns 31821 ns 1.04
array/reductions/mapreduce/Float32/dims=1 38163 ns 37330 ns 1.02
array/reductions/mapreduce/Float32/dims=1L 51893 ns 50978 ns 1.02
array/reductions/mapreduce/Float32/dims=2 56180 ns 54829 ns 1.02
array/reductions/mapreduce/Float32/dims=2L 68634 ns 67220 ns 1.02
array/reductions/mapreduce/Int64/1d 40006 ns 39331 ns 1.02
array/reductions/mapreduce/Int64/dims=1 41009 ns 40315 ns 1.02
array/reductions/mapreduce/Int64/dims=1L 89407 ns 88477 ns 1.01
array/reductions/mapreduce/Int64/dims=2 58019 ns 57574 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 84975 ns 83461 ns 1.02
array/reductions/reduce/Float32/1d 33410 ns 32196 ns 1.04
array/reductions/reduce/Float32/dims=1 38185 ns 37272 ns 1.02
array/reductions/reduce/Float32/dims=1L 51471 ns 50697 ns 1.02
array/reductions/reduce/Float32/dims=2 55858 ns 55042 ns 1.01
array/reductions/reduce/Float32/dims=2L 68605 ns 67551 ns 1.02
array/reductions/reduce/Int64/1d 40226 ns 39026 ns 1.03
array/reductions/reduce/Int64/dims=1 40773 ns 40114 ns 1.02
array/reductions/reduce/Int64/dims=1L 89248 ns 88494 ns 1.01
array/reductions/reduce/Int64/dims=2 58113 ns 57350 ns 1.01
array/reductions/reduce/Int64/dims=2L 84847 ns 83169 ns 1.02
array/reverse/1d 17414 ns 17038 ns 1.02
array/reverse/1dL 70259 ns 69729 ns 1.01
array/reverse/1dL_inplace 67769 ns 67315 ns 1.01
array/reverse/1d_inplace 10713 ns 8467.666666666666 ns 1.27
array/reverse/2d 20892 ns 20236 ns 1.03
array/reverse/2dL 74434 ns 73577 ns 1.01
array/reverse/2dL_inplace 67527 ns 67185 ns 1.01
array/reverse/2d_inplace 12265 ns 9814 ns 1.25
array/sorting/1d 2648173 ns 2646705 ns 1.00
array/sorting/2d 1019394 ns 1017870 ns 1.00
array/sorting/by 3175618 ns 3174206 ns 1.00
cuda/synchronization/context/auto 6782.6 ns 6849.6 ns 0.99
cuda/synchronization/context/blocking 789.8350515463917 ns 815.9772727272727 ns 0.97
cuda/synchronization/context/nonblocking 6873 ns 6897.4 ns 1.00
cuda/synchronization/stream/auto 684.8055555555555 ns 700.5918367346939 ns 0.98
cuda/synchronization/stream/blocking 863.0816326530612 ns 892.3673469387755 ns 0.97
cuda/synchronization/stream/nonblocking 7290.25 ns 7433 ns 0.98
integration/byval/reference 148662 ns 148108 ns 1.00
integration/byval/slices=1 149702 ns 149157 ns 1.00
integration/byval/slices=2 292290 ns 291918 ns 1.00
integration/byval/slices=3 435120 ns 434943 ns 1.00
integration/cudadevrt 105495 ns 105211 ns 1.00
integration/volumerhs 9154041 ns 9139091 ns 1.00
kernel/indexing 13427 ns 13078 ns 1.03
kernel/indexing_checked 14000 ns 13558 ns 1.03
kernel/launch 2425 ns 2086.222222222222 ns 1.16
kernel/occupancy 997.4615384615385 ns 684.6734693877551 ns 1.46
kernel/rand 16286 ns 16385 ns 0.99
latency/import 4277248384 ns 4265954564 ns 1.00
latency/precompile 5041012439 ns 5038240985 ns 1.00
latency/ttfp 4722095061 ns 4715207932 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Releasing wrapped memory last used on a special or destroyed stream waits for
the entire context, which is prohibited while any of its streams is being
captured, regardless of the capture mode. Postpone that until captures made
with `capture` have ended.
@maleadt
maleadt force-pushed the tb/wrap-release-capture branch from c1bd27b to f1faae0 Compare October 3, 2026 07:58
@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.50000% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 86.23%. Comparing base (a3202f2) to head (f1faae0).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
CUDACore/src/array.jl 81.25% 3 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##             main    #3320       +/-   ##
===========================================
+ Coverage   39.40%   86.23%   +46.83%     
===========================================
  Files         187      187               
  Lines       18966    18997       +31     
===========================================
+ Hits         7474    16383     +8909     
+ Misses      11492     2614     -8878     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt merged commit c28bedd into main Oct 3, 2026
1 check passed
@maleadt
maleadt deleted the tb/wrap-release-capture branch October 3, 2026 13:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant