Skip to content

Don't wait for the GPU when releasing resources from finalizers - #3318

Open
maleadt wants to merge 8 commits into
mainfrom
tb/deferred-release
Open

maleadt wants to merge 8 commits into
mainfrom
tb/deferred-release

Conversation

@maleadt

@maleadt maleadt commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Julia runs finalizers on whatever thread happens to trigger a garbage collection, at whatever point that thread allocates or releases a lock. CUDA.jl uses finalizers to free GPU memory and destroy other objects, but several of the driver and library calls it made there wait for all running kernels to finish, on any stream. Worse, while one thread is inside such a call, kernel launches from every other thread are blocked too. Measured with a 1 s kernel running on another stream, on an RTX 5080 (the Jetson Nano, Xavier and Orin behave the same for the memory calls):

call made from a finalizer blocks the collecting thread blocks a launch on another thread
cuMemFree (unified memory, or JULIA_CUDA_MEMORY_POOL=none) 906 ms 805 ms
cuMemFreeHost (HostMemory arrays) 906 ms 805 ms
cuMemHostUnregister (pin, unsafe_wrap of host memory) 906 ms 805 ms
cuArrayDestroy (texture arrays) 906 ms 805 ms
cudnnDestroy, curandDestroyGenerator, some cuSPARSE info objects 899 ms 799 ms
cuModuleUnload, cuBLAS/cuFFT/cuSOLVER handle destruction 899 ms –
cuMemFreeAsync (the default stream-ordered pool) – –

So a GC at the wrong moment stalls a thread, and every other thread's launches, for as long as unrelated kernels run. That includes the GC that synchronize may trigger right before it starts waiting. And if the running GPU work depends on a task on that thread making progress (e.g. a kernel waiting for a flag the host sets), it deadlocks. Moving the calls to a background thread doesn't help, because of the global launch stall: they have to be avoided while kernels run.

With this PR, finalizers never call into CUDA. They only push the object onto a lock-free list, and regular tasks release it later: when allocating, around synchronization, and otherwise within a second. What happens then depends on whether releasing the resource can wait for the GPU:

  • Releases that don't wait happen right away. Device memory is freed in stream order with cuMemFreeAsync, which now also covers JULIA_CUDA_MEMORY_POOL=none. Pinned host and unified memory is allocated from stream-ordered pools where the driver supports them (CUDA 13+), and freed the same way. Where it doesn't (e.g. on Jetson devices), freed host and unified allocations are cached and reused once the work using them has finished. Events, graphs, texture objects and library objects known not to wait are destroyed, and library handles go back to their cache.
  • Releases that may wait are held until memory is reclaimed, i.e. on an out-of-memory error or when calling CUDA.reclaim(): unregistering pinned or wrapped host memory (the owner is kept alive until then), emptying the host and unified memory caches, freeing memory without stream-ordered support (very old drivers), unloading modules, destroying texture arrays, and library objects whose destruction may wait (cuRAND generators, cuSPARSE analysis infos, cuTENSOR plans, ...). The background task that trims an idle pool doesn't reclaim anymore.

Memory is often released after the stream it was last used on has been collected, so a destroyed stream records a final event, and memory released afterwards waits for it on a separate non-blocking stream per context.

Every finalizer that CUDA.jl registers goes through a single function, resource_finalizer(f, obj; blocking=true, ctx=context()), which is now public API (CUDA.resource_finalizer). The finalizer only queues the object; f runs later on a regular task, in the context captured at registration: at the next drain if blocking=false, or when memory is reclaimed (the default). Packages wrapping CUDA library objects can use it for the same purpose.

No lock is held while collecting garbage (a task holding a lock doesn't run finalizers, so reclaim() wouldn't free what just became unreachable) or while waiting for the GPU (the work being waited for may depend on a task that needs that lock). Nothing is released while a graph is being captured, as that could invalidate the capture; starting a capture waits for routine releases to finish, and fails instead of waiting while reclaim() synchronizes the device or runs destructors that may wait for the GPU.

User-visible changes in behavior: collected memory isn't released by GC.gc() itself anymore, but by the next CUDA.jl operation that allocates or synchronizes, or within a second (CUDA.pool_status() releases it first). Resources whose release may wait for the GPU are released at the latest when CUDA.jl runs out of memory, or when calling CUDA.reclaim(), which can itself wait for the GPU. That includes the host and unified memory caches on devices without memory pools, which aren't trimmed in the background, as that would block. JULIA_CUDA_MEMORY_POOL=none now also disables the host and unified memory pools.

The design follows what other GPU stacks do: PyTorch, CuPy, XLA, RMM and CCCL don't release storage to the driver when the language object dies either, but recycle it in stream order, and only release it at explicit points like out-of-memory handling or empty_cache().

Not waiting for the GPU also makes allocating host and unified memory considerably faster, as pinning and unpinning, or allocating unified memory, is now mostly avoided. Allocating, touching and freeing an array in a loop, on the RTX 5080 (µs per iteration):

before after
DeviceMemory, 4 KiB 0.80 0.81
DeviceMemory, 4 KiB, JULIA_CUDA_MEMORY_POOL=none 32.7 1.7
HostMemory, 4 KiB 306 7.8
HostMemory, 64 MiB 13822 2988 (mostly the PCIe transfer)
UnifiedMemory, 4 KiB 29.2 1.7
UnifiedMemory, 64 MiB 7432 2407

The PR also contains a few independent fixes it depends on: kernel output is flushed before reporting a kernel exception (it used to be flushed as a side effect of blocking frees), a failed registration isn't recorded as a pin anymore, an event that synchronize found done by polling is synchronized once more (compute-sanitizer doesn't treat a successful cuEventQuery as synchronization, #3346), and a multi-GPU test restores the device it switched to.

The new tests keep the GPU busy with a kernel that waits for another task on the same thread, and then collect arrays of every memory type, pinned and wrapped arrays, texture arrays, modules and streams. Without this PR, these operations block until the kernel gives up after its 60 s timeout; with it, they pass in seconds. The resource, driver, array, pool and exception tests, and the library tests touched here, pass on an RTX 5080 with both memory allocators.

@maleadt
maleadt added this pull request to stack #3319 October 1, 2026 11:24
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 7b3fbb6 Previous: e924726 Ratio
array/accumulate/Float32/1d 98870 ns 99107 ns 1.00
array/accumulate/Float32/dims=1 73231 ns 72803 ns 1.01
array/accumulate/Float32/dims=1L 1589200 ns 1589176 ns 1.00
array/accumulate/Float32/dims=2 139011 ns 138542 ns 1.00
array/accumulate/Float32/dims=2L 656831 ns 655298 ns 1.00
array/accumulate/Int64/1d 117840 ns 117327 ns 1.00
array/accumulate/Int64/dims=1 77137 ns 77135 ns 1.00
array/accumulate/Int64/dims=1L 1710794 ns 1698128 ns 1.01
array/accumulate/Int64/dims=2 151533 ns 150131 ns 1.01
array/accumulate/Int64/dims=2L 987691 ns 986795 ns 1.00
array/broadcast 16365 ns 16029 ns 1.02
array/broadcast launch 7412.25 ns 7358.25 ns 1.01
array/construct 926.65625 ns 959.9 ns 0.97
array/copy 16228 ns 16576 ns 0.98
array/copyto!/cpu_to_gpu 206200 ns 207398 ns 0.99
array/copyto!/gpu_to_cpu 241228 ns 239800 ns 1.01
array/copyto!/gpu_to_gpu 10045 ns 8720.666666666666 ns 1.15
array/iteration/findall/bool 129777 ns 129599 ns 1.00
array/iteration/findall/int 140251 ns 140825 ns 1.00
array/iteration/findfirst/bool 67934 ns 67856 ns 1.00
array/iteration/findfirst/int 69901 ns 69203 ns 1.01
array/iteration/findmin/1d 65208 ns 59849 ns 1.09
array/iteration/findmin/2d 97744 ns 97040 ns 1.01
array/iteration/logical 182539 ns 181916 ns 1.00
array/iteration/scalar 60019 ns 58365 ns 1.03
array/permutedims/2d 46320 ns 45704 ns 1.01
array/permutedims/3d 47174 ns 47728 ns 0.99
array/permutedims/4d 49114 ns 48656 ns 1.01
array/random/rand/Float32 11285 ns 10851 ns 1.04
array/random/rand/Int64 19285 ns 18300 ns 1.05
array/random/rand!/Float32 7891.666666666667 ns 7842.666666666667 ns 1.01
array/random/rand!/Int64 17103 ns 16741 ns 1.02
array/random/randn/Float32 32776 ns 32264 ns 1.02
array/random/randn!/Float32 23418 ns 24124 ns 0.97
array/reductions/mapreduce/Float32/1d 32843 ns 32799 ns 1.00
array/reductions/mapreduce/Float32/dims=1 38098 ns 37628 ns 1.01
array/reductions/mapreduce/Float32/dims=1L 51665 ns 51393 ns 1.01
array/reductions/mapreduce/Float32/dims=2 55924 ns 55676 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 68057 ns 67890 ns 1.00
array/reductions/mapreduce/Int64/1d 39703 ns 39071 ns 1.02
array/reductions/mapreduce/Int64/dims=1 40941 ns 40979 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 89132 ns 89303 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57964 ns 57656 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 84418 ns 84117 ns 1.00
array/reductions/reduce/Float32/1d 33185 ns 32726 ns 1.01
array/reductions/reduce/Float32/dims=1 38049 ns 37671 ns 1.01
array/reductions/reduce/Float32/dims=1L 51481 ns 51035 ns 1.01
array/reductions/reduce/Float32/dims=2 55711 ns 55445 ns 1.00
array/reductions/reduce/Float32/dims=2L 68417 ns 68151 ns 1.00
array/reductions/reduce/Int64/1d 40163 ns 39207 ns 1.02
array/reductions/reduce/Int64/dims=1 40848 ns 40842 ns 1.00
array/reductions/reduce/Int64/dims=1L 89170 ns 88981 ns 1.00
array/reductions/reduce/Int64/dims=2 58097 ns 58095 ns 1.00
array/reductions/reduce/Int64/dims=2L 84448 ns 84484 ns 1.00
array/reverse/1d 17582 ns 17482 ns 1.01
array/reverse/1dL 70226 ns 70086 ns 1.00
array/reverse/1dL_inplace 67915 ns 67824 ns 1.00
array/reverse/1d_inplace 10678 ns 9051.333333333334 ns 1.18
array/reverse/2d 20746 ns 20350 ns 1.02
array/reverse/2dL 74015 ns 73763 ns 1.00
array/reverse/2dL_inplace 67698 ns 67393 ns 1.00
array/reverse/2d_inplace 12434 ns 10090 ns 1.23
array/sorting/1d 2646215 ns 2646314 ns 1.00
array/sorting/2d 1018365 ns 1018011 ns 1.00
array/sorting/by 3173138 ns 3158564 ns 1.00
cuda/synchronization/context/auto 7807.25 ns 6775.2 ns 1.15
cuda/synchronization/context/blocking 882.5625 ns 808.9775280898876 ns 1.09
cuda/synchronization/context/nonblocking 7836.75 ns 6745.2 ns 1.16
cuda/synchronization/stream/auto 735.888 ns 707.3802816901408 ns 1.04
cuda/synchronization/stream/blocking 953.1904761904761 ns 862.0714285714286 ns 1.11
cuda/synchronization/stream/nonblocking 8219.333333333334 ns 7079.5 ns 1.16
integration/byval/reference 148454 ns 148360 ns 1.00
integration/byval/slices=1 149518 ns 149203 ns 1.00
integration/byval/slices=2 292438 ns 292017 ns 1.00
integration/byval/slices=3 435624 ns 434881 ns 1.00
integration/cudadevrt 105841 ns 105387 ns 1.00
integration/volumerhs 9154132 ns 9145562 ns 1.00
kernel/indexing 13666 ns 13190 ns 1.04
kernel/indexing_checked 14236 ns 13943 ns 1.02
kernel/launch 2509.4444444444443 ns 2472 ns 1.02
kernel/occupancy 948.1666666666666 ns 938.1739130434783 ns 1.01
kernel/rand 17581 ns 14010 ns 1.25
latency/import 4335145735 ns 4302362850 ns 1.01
latency/precompile 5095975920 ns 5082487750 ns 1.00
latency/ttfp 4800382421 ns 4807353177 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Base automatically changed from tb/coopsync to main October 2, 2026 05:43
@maleadt
maleadt force-pushed the tb/deferred-release branch from 11a8eed to 55fdcc1 Compare October 3, 2026 14:16
@maleadt
maleadt removed this pull request from stack #3319 October 3, 2026 15:49
@maleadt
maleadt added this pull request to stack #3327 October 3, 2026 15:49
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.45050% with 61 lines in your changes missing coverage. Please review.
✅ Project coverage is 87.01%. Comparing base (e924726) to head (7b3fbb6).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
CUDACore/src/resources/pools.jl 45.65% 25 Missing ⚠️
CUDACore/src/resources/backends.jl 93.71% 12 Missing ⚠️
CUDACore/src/resources/pins.jl 91.00% 9 Missing ⚠️
CUDACore/src/resources/lifecycle.jl 95.74% 6 Missing ⚠️
CUDACore/src/resources/managed.jl 94.00% 3 Missing ⚠️
CUDACore/src/resources/streams.jl 95.00% 3 Missing ⚠️
lib/cusparse/src/helpers.jl 91.30% 2 Missing ⚠️
CUDACore/lib/cudadrv/graph.jl 83.33% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3318      +/-   ##
==========================================
+ Coverage   86.80%   87.01%   +0.20%     
==========================================
  Files         187      194       +7     
  Lines       18820    19201     +381     
==========================================
+ Hits        16337    16708     +371     
- Misses       2483     2493      +10     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt force-pushed the tb/deferred-release branch from 55fdcc1 to 5e2e801 Compare October 4, 2026 15:07
@maleadt maleadt changed the title Don't wait for the GPU when freeing memory from finalizers Don't wait for the GPU when releasing resources from finalizers Oct 4, 2026
@maleadt
maleadt force-pushed the tb/deferred-release branch 2 times, most recently from 9e797a7 to 3fdb466 Compare October 5, 2026 00:17
@maleadt
maleadt removed this pull request from stack #3327 October 5, 2026 08:33
@maleadt
maleadt added this pull request to stack #3338 October 5, 2026 08:33
@maleadt
maleadt force-pushed the tb/deferred-release branch 2 times, most recently from 9f24aa2 to 1115dac Compare October 6, 2026 08:24
compute-sanitizer doesn't treat a successful cuEventQuery as
synchronization (#3346), so it reports use-after-free errors for memory
that is released after waiting for an event, e.g., when releasing memory
last used on another stream. Synchronizing an event that has completed
doesn't block, unless it has been recorded again in the meantime, and
takes about 70 ns.
Separates the code that manages the lifetime of memory, i.e., the
memory pools, the Managed wrapper, pool_alloc and pool_free, reclaiming,
and pinning, from the rest of memory.jl. No functional changes.
The test left its worker on another device, which affected the tests
that ran on that worker later on.
Synchronizing in a non-blocking way polls for completion, which doesn't
flush the device's printf buffer. So the diagnostic output of a kernel
that threw an exception was only printed by a later blocking call, e.g.
when freeing memory or at exit, after the exception had been reported.
Synchronize an event that was never recorded, which flushes the buffer
without waiting for any work, also when another stream is still busy.
Relax the capture mode for that call, so that reporting an exception
from an earlier kernel doesn't invalidate a capture.
When registering failed, e.g. because the memory was registered already,
the pin count had been incremented and the object recorded, so pinning
the object again returned early without registering it. Also forget the
pin count when unpinning, and only register a finalizer the first time
an object is pinned: re-pinning a resized array registered another one,
which unpinned the memory a second time.
Finalizers run on whatever thread triggers a collection, and several of
the driver calls CUDA.jl made there (cuMemFree without a pool,
cuMemFreeHost, cuMemHostUnregister, cuArrayDestroy, cuModuleUnload)
wait for all running kernels to finish, also blocking kernel launches
from other threads meanwhile. If the running work depends on a task on
the collecting thread, that deadlocks.

Finalizers now only push the resource onto a lock-free list, and
regular tasks release it: when allocating (a bounded number of them),
around synchronization, and once a second. Device memory is freed in
stream order, also when it wasn't allocated from a pool. Releases that
may wait for the GPU (freeing host and unified memory, unregistering
host memory, freeing memory without stream-ordered support, unloading
modules, destroying texture arrays) are held until memory is reclaimed,
i.e., when running out of memory or calling reclaim(), so the task that
trims an idle pool doesn't reclaim anymore. Owners of wrapped host
memory are kept alive until the GPU has finished using it. Releases
that fail keep their resource alive instead of being retried, as they
may have partially succeeded.

Memory is often released after the stream it was last used on has been
collected, so streams record a final event before being destroyed, and
memory released afterwards waits for it on a non-blocking stream that
every context gets when it is first used. Memory last used on the
per-thread default stream, which can't be waited for from another
thread, is only released when reclaiming memory, after synchronizing the
context it was used in (which Managed now records).

Every finalizer in CUDACore is registered through resource_finalizer,
which is public so that packages wrapping library objects can do the
same. Handles returned to a HandleCache stay reusable until memory is
reclaimed, as destroying them may wait for the GPU.

No lock is held while collecting garbage, which keeps finalizers from
running, or while waiting for the GPU, as that work may depend on a task
that needs the lock. Releasing resources and starting a graph capture
exclude each other using short critical sections; a capture fails
instead of waiting while reclaim() synchronizes the device or releases
resources that may wait for the GPU.
Freeing host and unified memory waits for all running kernels to
finish, so it is only done when memory is reclaimed, which made
allocating such memory slow and its usage grow until running out.
Where supported (CUDA 13+), allocate it from stream-ordered memory
pools instead, which are freed in stream order like device memory.
Otherwise, e.g. on Jetson devices, cache freed allocations, rounded up
to a size class, and reuse them once an event recorded after their last
use has completed. Cached memory is released when memory is reclaimed.
Memory imported using unsafe_wrap(; own=true) is never reused, and is
freed with cuMemFreeAsync if it was allocated from a pool.

JULIA_CUDA_MEMORY_POOL=none also disables these pools.

Copies between pointers now use cuMemcpyAsync, as cuMemcpyDtoDAsync
reads memory from a host pool when the copy is enqueued instead of in
stream order. Memory from a host pool reports the device of the pool's
location, so only use a peer copy between actual device memory.
Register the finalizers of library handles, plans, descriptors and
helper objects with resource_finalizer, so they don't call into the
libraries from a finalizer. Destructors known not to wait for running
kernels (returning handles to their cache, cuDNN tensor descriptors,
cuSPARSE matrix and vector descriptors) run when retired resources are
released; the others (e.g. cuRAND generators, cuSOLVER Mg handles,
cuSPARSE analysis infos, cuTENSOR plans) when memory is reclaimed, in
the context they were created in.

cuTENSOR plans are now destroyed before releasing their workspace, and
reclaiming memory destroys the cuRAND generators in the cache it purges
instead of leaving them for a later collection. Drop a redundant
finalizer on the pointer array of the batched cuBLAS getrf wrapper.
@maleadt
maleadt force-pushed the tb/deferred-release branch from 1115dac to 7b3fbb6 Compare October 6, 2026 13:55

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant