Repository navigation
Conversation
maleadt
added this pull request to stack #3319
October 1, 2026 11:24
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
maleadt
force-pushed
the
tb/deferred-release
branch
from
October 3, 2026 14:16
11a8eed to
55fdcc1
Compare
maleadt
removed this pull request from stack #3319
October 3, 2026 15:49
maleadt
added this pull request to stack #3327
October 3, 2026 15:49
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #3318 +/- ##
==========================================
+ Coverage 86.80% 87.01% +0.20%
==========================================
Files 187 194 +7
Lines 18820 19201 +381
==========================================
+ Hits 16337 16708 +371
- Misses 2483 2493 +10 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
maleadt
force-pushed
the
tb/deferred-release
branch
from
October 4, 2026 15:07
55fdcc1 to
5e2e801
Compare
maleadt
force-pushed
the
tb/deferred-release
branch
2 times, most recently
from
October 5, 2026 00:17
9e797a7 to
3fdb466
Compare
maleadt
removed this pull request from stack #3327
October 5, 2026 08:33
maleadt
added this pull request to stack #3338
October 5, 2026 08:33
maleadt
force-pushed
the
tb/deferred-release
branch
2 times, most recently
from
October 6, 2026 08:24
9f24aa2 to
1115dac
Compare
compute-sanitizer doesn't treat a successful cuEventQuery as synchronization (#3346), so it reports use-after-free errors for memory that is released after waiting for an event, e.g., when releasing memory last used on another stream. Synchronizing an event that has completed doesn't block, unless it has been recorded again in the meantime, and takes about 70 ns.
Separates the code that manages the lifetime of memory, i.e., the memory pools, the Managed wrapper, pool_alloc and pool_free, reclaiming, and pinning, from the rest of memory.jl. No functional changes.
The test left its worker on another device, which affected the tests that ran on that worker later on.
Synchronizing in a non-blocking way polls for completion, which doesn't flush the device's printf buffer. So the diagnostic output of a kernel that threw an exception was only printed by a later blocking call, e.g. when freeing memory or at exit, after the exception had been reported. Synchronize an event that was never recorded, which flushes the buffer without waiting for any work, also when another stream is still busy. Relax the capture mode for that call, so that reporting an exception from an earlier kernel doesn't invalidate a capture.
When registering failed, e.g. because the memory was registered already, the pin count had been incremented and the object recorded, so pinning the object again returned early without registering it. Also forget the pin count when unpinning, and only register a finalizer the first time an object is pinned: re-pinning a resized array registered another one, which unpinned the memory a second time.
Finalizers run on whatever thread triggers a collection, and several of the driver calls CUDA.jl made there (cuMemFree without a pool, cuMemFreeHost, cuMemHostUnregister, cuArrayDestroy, cuModuleUnload) wait for all running kernels to finish, also blocking kernel launches from other threads meanwhile. If the running work depends on a task on the collecting thread, that deadlocks. Finalizers now only push the resource onto a lock-free list, and regular tasks release it: when allocating (a bounded number of them), around synchronization, and once a second. Device memory is freed in stream order, also when it wasn't allocated from a pool. Releases that may wait for the GPU (freeing host and unified memory, unregistering host memory, freeing memory without stream-ordered support, unloading modules, destroying texture arrays) are held until memory is reclaimed, i.e., when running out of memory or calling reclaim(), so the task that trims an idle pool doesn't reclaim anymore. Owners of wrapped host memory are kept alive until the GPU has finished using it. Releases that fail keep their resource alive instead of being retried, as they may have partially succeeded. Memory is often released after the stream it was last used on has been collected, so streams record a final event before being destroyed, and memory released afterwards waits for it on a non-blocking stream that every context gets when it is first used. Memory last used on the per-thread default stream, which can't be waited for from another thread, is only released when reclaiming memory, after synchronizing the context it was used in (which Managed now records). Every finalizer in CUDACore is registered through resource_finalizer, which is public so that packages wrapping library objects can do the same. Handles returned to a HandleCache stay reusable until memory is reclaimed, as destroying them may wait for the GPU. No lock is held while collecting garbage, which keeps finalizers from running, or while waiting for the GPU, as that work may depend on a task that needs the lock. Releasing resources and starting a graph capture exclude each other using short critical sections; a capture fails instead of waiting while reclaim() synchronizes the device or releases resources that may wait for the GPU.
Freeing host and unified memory waits for all running kernels to finish, so it is only done when memory is reclaimed, which made allocating such memory slow and its usage grow until running out. Where supported (CUDA 13+), allocate it from stream-ordered memory pools instead, which are freed in stream order like device memory. Otherwise, e.g. on Jetson devices, cache freed allocations, rounded up to a size class, and reuse them once an event recorded after their last use has completed. Cached memory is released when memory is reclaimed. Memory imported using unsafe_wrap(; own=true) is never reused, and is freed with cuMemFreeAsync if it was allocated from a pool. JULIA_CUDA_MEMORY_POOL=none also disables these pools. Copies between pointers now use cuMemcpyAsync, as cuMemcpyDtoDAsync reads memory from a host pool when the copy is enqueued instead of in stream order. Memory from a host pool reports the device of the pool's location, so only use a peer copy between actual device memory.
Register the finalizers of library handles, plans, descriptors and helper objects with resource_finalizer, so they don't call into the libraries from a finalizer. Destructors known not to wait for running kernels (returning handles to their cache, cuDNN tensor descriptors, cuSPARSE matrix and vector descriptors) run when retired resources are released; the others (e.g. cuRAND generators, cuSOLVER Mg handles, cuSPARSE analysis infos, cuTENSOR plans) when memory is reclaimed, in the context they were created in. cuTENSOR plans are now destroyed before releasing their workspace, and reclaiming memory destroys the cuRAND generators in the cache it purges instead of leaving them for a later collection. Drop a redundant finalizer on the pointer array of the batched cuBLAS getrf wrapper.
maleadt
force-pushed
the
tb/deferred-release
branch
from
October 6, 2026 13:55
1115dac to
7b3fbb6
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Julia runs finalizers on whatever thread happens to trigger a garbage collection, at whatever point that thread allocates or releases a lock. CUDA.jl uses finalizers to free GPU memory and destroy other objects, but several of the driver and library calls it made there wait for all running kernels to finish, on any stream. Worse, while one thread is inside such a call, kernel launches from every other thread are blocked too. Measured with a 1 s kernel running on another stream, on an RTX 5080 (the Jetson Nano, Xavier and Orin behave the same for the memory calls):
cuMemFree(unified memory, orJULIA_CUDA_MEMORY_POOL=none)cuMemFreeHost(HostMemoryarrays)cuMemHostUnregister(pin,unsafe_wrapof host memory)cuArrayDestroy(texture arrays)cudnnDestroy,curandDestroyGenerator, some cuSPARSE info objectscuModuleUnload, cuBLAS/cuFFT/cuSOLVER handle destructioncuMemFreeAsync(the default stream-ordered pool)So a GC at the wrong moment stalls a thread, and every other thread's launches, for as long as unrelated kernels run. That includes the GC that
synchronizemay trigger right before it starts waiting. And if the running GPU work depends on a task on that thread making progress (e.g. a kernel waiting for a flag the host sets), it deadlocks. Moving the calls to a background thread doesn't help, because of the global launch stall: they have to be avoided while kernels run.With this PR, finalizers never call into CUDA. They only push the object onto a lock-free list, and regular tasks release it later: when allocating, around synchronization, and otherwise within a second. What happens then depends on whether releasing the resource can wait for the GPU:
cuMemFreeAsync, which now also coversJULIA_CUDA_MEMORY_POOL=none. Pinned host and unified memory is allocated from stream-ordered pools where the driver supports them (CUDA 13+), and freed the same way. Where it doesn't (e.g. on Jetson devices), freed host and unified allocations are cached and reused once the work using them has finished. Events, graphs, texture objects and library objects known not to wait are destroyed, and library handles go back to their cache.CUDA.reclaim(): unregistering pinned or wrapped host memory (the owner is kept alive until then), emptying the host and unified memory caches, freeing memory without stream-ordered support (very old drivers), unloading modules, destroying texture arrays, and library objects whose destruction may wait (cuRAND generators, cuSPARSE analysis infos, cuTENSOR plans, ...). The background task that trims an idle pool doesn't reclaim anymore.Memory is often released after the stream it was last used on has been collected, so a destroyed stream records a final event, and memory released afterwards waits for it on a separate non-blocking stream per context.
Every finalizer that CUDA.jl registers goes through a single function,
resource_finalizer(f, obj; blocking=true, ctx=context()), which is now public API (CUDA.resource_finalizer). The finalizer only queues the object;fruns later on a regular task, in the context captured at registration: at the next drain ifblocking=false, or when memory is reclaimed (the default). Packages wrapping CUDA library objects can use it for the same purpose.No lock is held while collecting garbage (a task holding a lock doesn't run finalizers, so
reclaim()wouldn't free what just became unreachable) or while waiting for the GPU (the work being waited for may depend on a task that needs that lock). Nothing is released while a graph is being captured, as that could invalidate the capture; starting a capture waits for routine releases to finish, and fails instead of waiting whilereclaim()synchronizes the device or runs destructors that may wait for the GPU.User-visible changes in behavior: collected memory isn't released by
GC.gc()itself anymore, but by the next CUDA.jl operation that allocates or synchronizes, or within a second (CUDA.pool_status()releases it first). Resources whose release may wait for the GPU are released at the latest when CUDA.jl runs out of memory, or when callingCUDA.reclaim(), which can itself wait for the GPU. That includes the host and unified memory caches on devices without memory pools, which aren't trimmed in the background, as that would block.JULIA_CUDA_MEMORY_POOL=nonenow also disables the host and unified memory pools.The design follows what other GPU stacks do: PyTorch, CuPy, XLA, RMM and CCCL don't release storage to the driver when the language object dies either, but recycle it in stream order, and only release it at explicit points like out-of-memory handling or
empty_cache().Not waiting for the GPU also makes allocating host and unified memory considerably faster, as pinning and unpinning, or allocating unified memory, is now mostly avoided. Allocating, touching and freeing an array in a loop, on the RTX 5080 (µs per iteration):
DeviceMemory, 4 KiBDeviceMemory, 4 KiB,JULIA_CUDA_MEMORY_POOL=noneHostMemory, 4 KiBHostMemory, 64 MiBUnifiedMemory, 4 KiBUnifiedMemory, 64 MiBThe PR also contains a few independent fixes it depends on: kernel output is flushed before reporting a kernel exception (it used to be flushed as a side effect of blocking frees), a failed registration isn't recorded as a pin anymore, an event that
synchronizefound done by polling is synchronized once more (compute-sanitizer doesn't treat a successfulcuEventQueryas synchronization, #3346), and a multi-GPU test restores the device it switched to.The new tests keep the GPU busy with a kernel that waits for another task on the same thread, and then collect arrays of every memory type, pinned and wrapped arrays, texture arrays, modules and streams. Without this PR, these operations block until the kernel gives up after its 60 s timeout; with it, they pass in seconds. The resource, driver, array, pool and exception tests, and the library tests touched here, pass on an RTX 5080 with both memory allocators.