Skip to content

Recycle the streams of finished tasks - #3317

Open
maleadt wants to merge 4 commits into
tb/deferred-releasefrom
tb/stream-pool
Open

maleadt wants to merge 4 commits into
tb/deferred-releasefrom
tb/stream-pool

Conversation

@maleadt

@maleadt maleadt commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

CUDA.jl gives each Julia task its own stream. That keeps concurrently running tasks independent, but a finished task's stream stays alive until the GC collects the task. Each stream holds on to about half a MiB of device memory, and the stream-ordered allocator gets slower with every stream that has used it, so applications that spawn many short GPU tasks accumulate thousands of streams and slow down allocation for every task.

On an RTX 5080 (driver 610), 20,000 tasks that each compute sum(x .+ 1f0) on a small array took 17–35 s and reached 11,700 live streams (6.5 GiB of device memory). A producer/consumer workload with 20,000 iterations took 10.0 s with up to 15,300 streams and 9.2 GiB of device memory. The same work in a single task takes 0.71 s and 0.3 s respectively.

This PR keeps the streams of tasks in a pool per context, priority and flags, and hands a stream to the next task that needs one once its owner has finished and its GPU work has completed. Tasks running at the same time never share a stream, and up to 32 idle streams are kept around; idle streams beyond that are dropped the next time the pool is searched. With this, the two workloads above take 0.95 s with 2 streams (~350 MiB) and 0.45 s with 8 streams (~450 MiB).

Handing a stream to another task bumps a generation counter on the stream. Memory remembers the generation it was last used in, so after its stream has been recycled it knows that work has finished: using it doesn't wait for the new owner's work, and freeing it happens on the context's disposal stream instead of the recycled one, which the new owner may be capturing.

Because streams are now selected by the pool, this PR also adds task-level stream priorities. CUDA fixes a stream's priority when it is created, while Julia code schedules work in tasks, so priority! selects a stream of the requested priority for the current task and orders it after the task's previous stream:

CUDA.priority!(:high)            # subsequent GPU work in this task
# launch latency-sensitive work
CUDA.priority!(:normal)

CUDA.priority!(:high) do
    # work in this block uses high priority
end                              # restores the previous stream

priority! also accepts an integer from priority_range(), and KernelAbstractions.priority! uses the same mechanism (it used to create a new stream on every call). Selecting a priority before the task first uses the GPU creates only the requested stream, and switching back reuses the task's earlier stream. As in CUDA itself, priority is a scheduling hint for pending work, not preemption. Child tasks start at normal priority, and changing priority during graph capture is rejected. Using an array after a switch may wait on the CPU for work on its previous stream. If the block passed to priority! throws, the previous stream is selected again without ordering it after the block's work, so that the block's error is what gets reported.

The trade-off is that a stream returned by stream() belongs to the current task: code that keeps using it after the task has finished may share it with another task. Code that needs a stream beyond a task's lifetime should create one with CuStream(); the multitasking docs and NEWS say so. Looking for an idle stream scans the pool under a global lock, which is linear in the number of live tasks, but only happens when a task first uses the GPU or switches priority. The 32-stream retention limit is a fixed policy rather than a preference, since it only bounds idle streams.

Two small fixes come along: capture no longer ends the capture in progress when starting a nested one fails, and the docstring of priority(::CuStream) shows the right signature.

Tested on an RTX 5080 with the core/initialization, core/resources, core/cudadrv and core/kernelabstractions tests, including new tests that check that concurrent tasks get distinct streams, that the stream count stays bounded across 1,024 tasks, and that memory last used on a recycled stream is usable and freed without waiting for the stream's new owner.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 7623dd4 Previous: e924726 Ratio
array/accumulate/Float32/1d 98693 ns 99107 ns 1.00
array/accumulate/Float32/dims=1 72889 ns 72803 ns 1.00
array/accumulate/Float32/dims=1L 1588324 ns 1589176 ns 1.00
array/accumulate/Float32/dims=2 138705 ns 138542 ns 1.00
array/accumulate/Float32/dims=2L 655666 ns 655298 ns 1.00
array/accumulate/Int64/1d 118143 ns 117327 ns 1.01
array/accumulate/Int64/dims=1 77108 ns 77135 ns 1.00
array/accumulate/Int64/dims=1L 1697458 ns 1698128 ns 1.00
array/accumulate/Int64/dims=2 150384 ns 150131 ns 1.00
array/accumulate/Int64/dims=2L 987600 ns 986795 ns 1.00
array/broadcast 16115 ns 16029 ns 1.01
array/broadcast launch 7357.75 ns 7358.25 ns 1.00
array/construct 908.8648648648649 ns 959.9 ns 0.95
array/copy 16322 ns 16576 ns 0.98
array/copyto!/cpu_to_gpu 206559 ns 207398 ns 1.00
array/copyto!/gpu_to_cpu 240385 ns 239800 ns 1.00
array/copyto!/gpu_to_gpu 10187.333333333334 ns 8720.666666666666 ns 1.17
array/iteration/findall/bool 133581 ns 129599 ns 1.03
array/iteration/findall/int 145201 ns 140825 ns 1.03
array/iteration/findfirst/bool 70816 ns 67856 ns 1.04
array/iteration/findfirst/int 73133 ns 69203 ns 1.06
array/iteration/findmin/1d 69612 ns 59849 ns 1.16
array/iteration/findmin/2d 97571 ns 97040 ns 1.01
array/iteration/logical 189798 ns 181916 ns 1.04
array/iteration/scalar 55804 ns 58365 ns 0.96
array/permutedims/2d 46038 ns 45704 ns 1.01
array/permutedims/3d 47296 ns 47728 ns 0.99
array/permutedims/4d 48721 ns 48656 ns 1.00
array/random/rand/Float32 11542 ns 10851 ns 1.06
array/random/rand/Int64 19136 ns 18300 ns 1.05
array/random/rand!/Float32 7839 ns 7842.666666666667 ns 1.00
array/random/rand!/Int64 16767 ns 16741 ns 1.00
array/random/randn/Float32 32183 ns 32264 ns 1.00
array/random/randn!/Float32 24151 ns 24124 ns 1.00
array/reductions/mapreduce/Float32/1d 35199 ns 32799 ns 1.07
array/reductions/mapreduce/Float32/dims=1 38048 ns 37628 ns 1.01
array/reductions/mapreduce/Float32/dims=1L 51526 ns 51393 ns 1.00
array/reductions/mapreduce/Float32/dims=2 55485 ns 55676 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 67850 ns 67890 ns 1.00
array/reductions/mapreduce/Int64/1d 40423 ns 39071 ns 1.03
array/reductions/mapreduce/Int64/dims=1 40459 ns 40979 ns 0.99
array/reductions/mapreduce/Int64/dims=1L 89154 ns 89303 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57564 ns 57656 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 84176 ns 84117 ns 1.00
array/reductions/reduce/Float32/1d 34561 ns 32726 ns 1.06
array/reductions/reduce/Float32/dims=1 37840 ns 37671 ns 1.00
array/reductions/reduce/Float32/dims=1L 51407 ns 51035 ns 1.01
array/reductions/reduce/Float32/dims=2 55644 ns 55445 ns 1.00
array/reductions/reduce/Float32/dims=2L 68050 ns 68151 ns 1.00
array/reductions/reduce/Int64/1d 40843 ns 39207 ns 1.04
array/reductions/reduce/Int64/dims=1 40444 ns 40842 ns 0.99
array/reductions/reduce/Int64/dims=1L 89071 ns 88981 ns 1.00
array/reductions/reduce/Int64/dims=2 57843 ns 58095 ns 1.00
array/reductions/reduce/Int64/dims=2L 84012 ns 84484 ns 0.99
array/reverse/1d 17373 ns 17482 ns 0.99
array/reverse/1dL 70142 ns 70086 ns 1.00
array/reverse/1dL_inplace 67711 ns 67824 ns 1.00
array/reverse/1d_inplace 8830 ns 9051.333333333334 ns 0.98
array/reverse/2d 20682 ns 20350 ns 1.02
array/reverse/2dL 73887 ns 73763 ns 1.00
array/reverse/2dL_inplace 67589 ns 67393 ns 1.00
array/reverse/2d_inplace 10134 ns 10090 ns 1.00
array/sorting/1d 2656860 ns 2646314 ns 1.00
array/sorting/2d 1018349 ns 1018011 ns 1.00
array/sorting/by 3173032 ns 3158564 ns 1.00
cuda/synchronization/context/auto 7927 ns 6775.2 ns 1.17
cuda/synchronization/context/blocking 870.9107142857143 ns 808.9775280898876 ns 1.08
cuda/synchronization/context/nonblocking 7968.666666666667 ns 6745.2 ns 1.18
cuda/synchronization/stream/auto 730.3664122137404 ns 707.3802816901408 ns 1.03
cuda/synchronization/stream/blocking 956 ns 862.0714285714286 ns 1.11
cuda/synchronization/stream/nonblocking 8327 ns 7079.5 ns 1.18
integration/byval/reference 148318 ns 148360 ns 1.00
integration/byval/slices=1 149526 ns 149203 ns 1.00
integration/byval/slices=2 292142 ns 292017 ns 1.00
integration/byval/slices=3 435362 ns 434881 ns 1.00
integration/cudadevrt 105399 ns 105387 ns 1.00
integration/volumerhs 9152246 ns 9145562 ns 1.00
kernel/indexing 13299 ns 13190 ns 1.01
kernel/indexing_checked 14014 ns 13943 ns 1.01
kernel/launch 2369.1111111111113 ns 2472 ns 0.96
kernel/occupancy 924.2941176470588 ns 938.1739130434783 ns 0.99
kernel/rand 14141 ns 14010 ns 1.01
latency/import 4343834918 ns 4302362850 ns 1.01
latency/precompile 5092200784 ns 5082487750 ns 1.00
latency/ttfp 4850149145 ns 4807353177 ns 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.57143% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 87.09%. Comparing base (7b3fbb6) to head (7623dd4).

Files with missing lines Patch % Lines
CUDACore/lib/cudadrv/state.jl 98.14% 2 Missing ⚠️
Additional details and impacted files
@@                   Coverage Diff                   @@
##           tb/deferred-release    #3317      +/-   ##
=======================================================
+ Coverage                87.01%   87.09%   +0.07%     
=======================================================
  Files                      194      194              
  Lines                    19201    19310     +109     
=======================================================
+ Hits                     16708    16818     +110     
+ Misses                    2493     2492       -1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread CUDACore/lib/cudadrv/state.jl Outdated
Comment thread docs/src/usage/multitasking.md Outdated
@maleadt

maleadt commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

While working on this, I noticed AMDGPU.jl has a priority! function to change the priority of the current task. We didn't have that, instead requiring people to manually configure the underlying stream. With streams being recycled, that became a little iffy, so I had Sol port the task priority capability from AMDGPU.jl.

@maleadt
maleadt changed the base branch from main to tb/wrap-release-capture October 2, 2026 13:34
@maleadt
maleadt added this pull request to stack #3321 October 2, 2026 13:34
Base automatically changed from tb/wrap-release-capture to main October 3, 2026 13:11
@maleadt
maleadt removed this pull request from stack #3321 October 3, 2026 15:49
@maleadt
maleadt changed the base branch from main to tb/deferred-release October 3, 2026 15:49
@maleadt
maleadt added this pull request to stack #3327 October 3, 2026 15:49
@maleadt
maleadt removed this pull request from stack #3327 October 5, 2026 08:33
@maleadt
maleadt added this pull request to stack #3338 October 5, 2026 08:33
When cuStreamBeginCapture failed, e.g. because the stream was already
being captured, `capture` still called cuStreamEndCapture on the stream.
That ended the capture that was in progress, so its own `capture` call
failed later on, and the original error could be replaced by the one
from ending the capture.
Every task got its own stream, created on first use and only destroyed
when the GC finalized it. Each stream holds on to about half a MiB of
device memory, and the stream-ordered allocator gets slower with every
stream that has used it (an allocation takes ~1ms with 8000 such streams
alive). Since the GC is in no hurry to collect finished tasks, code that
spawns many short GPU tasks piles up thousands of streams.

Instead, keep the streams of tasks in a pool per context, and hand the
stream of a task that has finished, and whose work has completed, to the
next task that needs one. Tasks running at the same time never share a
stream, and up to 32 idle streams are kept around.

Handing a stream to another task bumps its generation, so that memory
last used by a recycled stream knows that its work has finished. Such
memory doesn't wait for the stream's new owner, and isn't freed on that
stream either, since the new owner may be capturing it.

Checking whether a stream is idle is prohibited while another thread
captures in global mode, which is when a newly spawned task may look
for a stream, so these queries use the relaxed capture mode.

Move the gate kernel from the cudadrv tests to the shared test helpers,
to keep a stream busy in the tests.
CUDA fixes the priority of a stream when it is created, while Julia code
schedules GPU work in tasks that each have their own stream. The only
way to change a task's priority was KernelAbstractions.priority!, which
created a new stream on every call and left it to the GC.

Add `priority!(p)`, which switches the current task to a stream of the
requested priority, taken from the stream pool (keyed by priority and
flags too now), and orders it after the task's previous stream.
Switching back reuses the stream the task used before, and selecting a
priority before the task first uses the GPU only creates the requested
stream. The do-block form restores the previous stream afterwards. If
the block throws, the stream is restored without ordering it after the
block's work, so that a failure to do so can't hide the original error.

Changing priority during graph capture is rejected, as it would move
the task's work out of the graph; `capture` marks the capturing task so
that this also covers streams selected with `stream!`.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants