Repository navigation
Conversation
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/deferred-release #3317 +/- ##
=======================================================
+ Coverage 87.01% 87.09% +0.07%
=======================================================
Files 194 194
Lines 19201 19310 +109
=======================================================
+ Hits 16708 16818 +110
+ Misses 2493 2492 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
kshyatt
reviewed
Oct 1, 2026
kshyatt
reviewed
Oct 1, 2026
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 1, 2026 18:04
20aceab to
9e6d20c
Compare
Member
Author
|
While working on this, I noticed AMDGPU.jl has a |
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 2, 2026 13:34
a77c408 to
51f3748
Compare
maleadt
added this pull request to stack #3321
October 2, 2026 13:34
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 3, 2026 07:58
51f3748 to
6cf803a
Compare
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 3, 2026 15:48
6cf803a to
e3226fd
Compare
maleadt
removed this pull request from stack #3321
October 3, 2026 15:49
maleadt
added this pull request to stack #3327
October 3, 2026 15:49
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 4, 2026 15:07
e3226fd to
6601a8c
Compare
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 4, 2026 19:19
6601a8c to
67f657a
Compare
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 5, 2026 00:17
67f657a to
b70ef84
Compare
maleadt
removed this pull request from stack #3327
October 5, 2026 08:33
maleadt
added this pull request to stack #3338
October 5, 2026 08:33
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 5, 2026 18:16
b70ef84 to
266c33e
Compare
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 6, 2026 08:24
266c33e to
c426449
Compare
When cuStreamBeginCapture failed, e.g. because the stream was already being captured, `capture` still called cuStreamEndCapture on the stream. That ended the capture that was in progress, so its own `capture` call failed later on, and the original error could be replaced by the one from ending the capture.
Every task got its own stream, created on first use and only destroyed when the GC finalized it. Each stream holds on to about half a MiB of device memory, and the stream-ordered allocator gets slower with every stream that has used it (an allocation takes ~1ms with 8000 such streams alive). Since the GC is in no hurry to collect finished tasks, code that spawns many short GPU tasks piles up thousands of streams. Instead, keep the streams of tasks in a pool per context, and hand the stream of a task that has finished, and whose work has completed, to the next task that needs one. Tasks running at the same time never share a stream, and up to 32 idle streams are kept around. Handing a stream to another task bumps its generation, so that memory last used by a recycled stream knows that its work has finished. Such memory doesn't wait for the stream's new owner, and isn't freed on that stream either, since the new owner may be capturing it. Checking whether a stream is idle is prohibited while another thread captures in global mode, which is when a newly spawned task may look for a stream, so these queries use the relaxed capture mode. Move the gate kernel from the cudadrv tests to the shared test helpers, to keep a stream busy in the tests.
CUDA fixes the priority of a stream when it is created, while Julia code schedules GPU work in tasks that each have their own stream. The only way to change a task's priority was KernelAbstractions.priority!, which created a new stream on every call and left it to the GC. Add `priority!(p)`, which switches the current task to a stream of the requested priority, taken from the stream pool (keyed by priority and flags too now), and orders it after the task's previous stream. Switching back reuses the stream the task used before, and selecting a priority before the task first uses the GPU only creates the requested stream. The do-block form restores the previous stream afterwards. If the block throws, the stream is restored without ordering it after the block's work, so that a failure to do so can't hide the original error. Changing priority during graph capture is rejected, as it would move the task's work out of the graph; `capture` marks the capturing task so that this also covers streams selected with `stream!`.
maleadt
force-pushed
the
tb/stream-pool
branch
from
October 6, 2026 13:55
c426449 to
7623dd4
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CUDA.jl gives each Julia task its own stream. That keeps concurrently running tasks independent, but a finished task's stream stays alive until the GC collects the task. Each stream holds on to about half a MiB of device memory, and the stream-ordered allocator gets slower with every stream that has used it, so applications that spawn many short GPU tasks accumulate thousands of streams and slow down allocation for every task.
On an RTX 5080 (driver 610), 20,000 tasks that each compute
sum(x .+ 1f0)on a small array took 17–35 s and reached 11,700 live streams (6.5 GiB of device memory). A producer/consumer workload with 20,000 iterations took 10.0 s with up to 15,300 streams and 9.2 GiB of device memory. The same work in a single task takes 0.71 s and 0.3 s respectively.This PR keeps the streams of tasks in a pool per context, priority and flags, and hands a stream to the next task that needs one once its owner has finished and its GPU work has completed. Tasks running at the same time never share a stream, and up to 32 idle streams are kept around; idle streams beyond that are dropped the next time the pool is searched. With this, the two workloads above take 0.95 s with 2 streams (~350 MiB) and 0.45 s with 8 streams (~450 MiB).
Handing a stream to another task bumps a generation counter on the stream. Memory remembers the generation it was last used in, so after its stream has been recycled it knows that work has finished: using it doesn't wait for the new owner's work, and freeing it happens on the context's disposal stream instead of the recycled one, which the new owner may be capturing.
Because streams are now selected by the pool, this PR also adds task-level stream priorities. CUDA fixes a stream's priority when it is created, while Julia code schedules work in tasks, so
priority!selects a stream of the requested priority for the current task and orders it after the task's previous stream:priority!also accepts an integer frompriority_range(), andKernelAbstractions.priority!uses the same mechanism (it used to create a new stream on every call). Selecting a priority before the task first uses the GPU creates only the requested stream, and switching back reuses the task's earlier stream. As in CUDA itself, priority is a scheduling hint for pending work, not preemption. Child tasks start at normal priority, and changing priority during graph capture is rejected. Using an array after a switch may wait on the CPU for work on its previous stream. If the block passed topriority!throws, the previous stream is selected again without ordering it after the block's work, so that the block's error is what gets reported.The trade-off is that a stream returned by
stream()belongs to the current task: code that keeps using it after the task has finished may share it with another task. Code that needs a stream beyond a task's lifetime should create one withCuStream(); the multitasking docs and NEWS say so. Looking for an idle stream scans the pool under a global lock, which is linear in the number of live tasks, but only happens when a task first uses the GPU or switches priority. The 32-stream retention limit is a fixed policy rather than a preference, since it only bounds idle streams.Two small fixes come along:
captureno longer ends the capture in progress when starting a nested one fails, and the docstring ofpriority(::CuStream)shows the right signature.Tested on an RTX 5080 with the
core/initialization,core/resources,core/cudadrvandcore/kernelabstractionstests, including new tests that check that concurrent tasks get distinct streams, that the stream count stays bounded across 1,024 tasks, and that memory last used on a recycled stream is usable and freed without waiting for the stream's new owner.