Repository navigation
Conversation
|
What happens under task parallelism? |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/stream-pool #3326 +/- ##
==================================================
+ Coverage 87.09% 87.15% +0.06%
==================================================
Files 194 195 +1
Lines 19310 19749 +439
==================================================
+ Hits 16818 17213 +395
- Misses 2492 2536 +44 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
🤖 Good question; it wasn't handled well. Concurrent captures, launching a graph from several tasks, and unrelated tasks using the GPU on other threads during a capture all worked, but:
The first two are fixed here: captures now use the relaxed mode and a dedicated stream, and kernel loading no longer synchronizes the device while capturing. Capturing the operations of spawned tasks, e.g. from libraries that parallelize internally, is #3337: they inherit the capture through a scoped value and submit to the same stream. What remains is a cuRAND issue (#3336): seeding a generator synchronizes the device, which breaks captures elsewhere. |
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
Memory can be in use beyond the lifetime of the object that owns it, e.g., by a graph that captured operations on an array. Add a lease count to `Managed` memory, and route every way of releasing managed memory (the pool, and wrapped host memory) through `release`, which postpones the release until the last lease has ended.
Capturing used to disable the GC for the whole process, as freeing memory while capturing breaks the capture. With deferred release, finalizers don't call into CUDA anymore, so instead of disabling the GC, don't dispose of retired resources while a capture is in progress, and drain them when it ends. Captures start under the drain lock, so they can't race with disposal that's already in progress. The graph functionality moves out of the driver wrappers, and captures now: - use the relaxed capture mode by default. The driver's checks for unsafe API calls apply per thread, not per task, so in the global or thread-local mode they also rejected operations of unrelated tasks that ran on the capturing thread while the capturing task was waiting, invalidating the capture. CUDA.jl checks the operations it performs itself, and the stricter modes remain available; - keep the capturing task on its thread in those stricter modes, as such captures need to end on the thread they started on; - end the capture on every exit path, and report the original error; - report waiting for the GPU while capturing with a descriptive `CaptureError`.
Executable graphs are now callable, and can be uploaded to the device ahead of their first launch. Graphs show their contents, can be rendered with Graphviz, and their nodes can be queried.
Graphs used to keep none of the memory their operations use alive, so launching a graph after one of its arrays had been freed accessed freed memory. Now, operations that use memory while capturing record that memory instead of synchronizing with its owner, and the resulting graph, as well as every executable graph instantiated from it, holds a lease on it. Launching an executable graph takes ownership of that memory like launching a kernel does, so that other tasks synchronize with it, and so that releasing the memory is ordered after the last launch. Arrays allocated from a `GPUArrays.@cached` scope are rejected while capturing, as the cache reuses their memory regardless of what still uses it.
Memory allocated in stream order while capturing becomes owned by the graph, valid only between launches, which doesn't fit arrays whose lifetime is managed by the GC (#1230). Instead, allocate such memory from the pool on a separate stream, and keep it alive using leases. That stream owns the memory, so whoever uses it next (launching the graph, or another task) waits for the allocation through the usual ownership mechanism. Every launch of the graph then overwrites the arrays that were allocated during capture. Scalar references (`CuRef`) are initialized right away instead of being captured, as memory copies are relatively expensive to launch as part of a graph. This also fixes capturing library calls with scalar arguments (#2691).
A captured copy from host memory reads that memory every time the graph is launched, long after the memory may have been freed. Copy such data into pinned memory that the graph keeps alive, giving captured copies snapshot semantics. Copying to host memory can't be made safe the same way, so it results in a `CaptureError`; arrays backed by `HostMemory` can be used instead.
Libraries may allocate memory in stream order while being captured, which becomes owned by the graph. Launching such a graph again requires freeing that memory first, so instantiate graphs with `AUTO_FREE_ON_LAUNCH`, and free the outstanding allocations when updating or destroying an executable graph, as the driver doesn't. Reclaiming memory now also trims the memory that graphs keep cached.
`capture!` adds captured operations to an existing graph, depending on given nodes, which makes it possible to construct graphs with operations that execute concurrently without juggling streams and events. `update!` changes the parameters of a single operation in an executable graph, e.g., to launch a kernel with different arguments, without capturing the graph again.
Capturing on the task's stream made it impossible for other tasks to wait for work the capturing task had submitted before the capture, e.g., to use an array it produced. Instead, check out a non-blocking stream that's only used for capturing, and make it the task's stream for the duration of the capture.
capture and capture! now accept a stream keyword argument, to capture on that stream instead of on a dedicated one. It needs to be a non-blocking stream of the current context, which is reserved for the capture, and is the task's stream while capturing.
The workaround for module loading under memory pressure synchronized the context unless the current task was capturing, which invalidated captures by other tasks. Skip it whenever a graph is being captured. Captures that start during the synchronization fail instead of waiting for it, as they do while reclaiming memory, since the awaited work may depend on the capturing task.
CUDA graphs let you record a sequence of GPU operations once and then launch the whole sequence with a single, cheap API call. For code made of many short kernels, which is common in Julia because every broadcast is a kernel, the CPU cost of launching each operation often exceeds the time the GPU spends executing it, so graphs can make such code several times faster. CUDA.jl has had
capture,instantiateand@capturedfor a while, but they didn't work with CUDA.jl's GC-based memory management:capturedisabled the GC for the whole process, and freeing memory from another thread could still invalidate the capture (Freeing memory from finalizers during graph capture #3310).This PR reworks the integration so that graphs are safe to use with ordinary Julia code, and adds an API for both simple capture-and-replay and for constructing graphs by hand. It builds on the deferred release of resources from #3318, which is what makes capture compatible with the GC. The commits are self-contained steps, and are best reviewed one at a time.
Allocating code can be captured as well, e.g.,
y .= W * x .+ b, whereW * xallocates a temporary: memory that is allocated during capture is allocated once, outside of the graph, kept alive by it, and overwritten by every launch.For building graphs by hand,
capture!adds captured operations to an existing graph with explicit dependencies, which makes it possible to express concurrency without juggling streams and events.update!changes the parameters of a single operation in an executable graph, e.g., to launch a kernel on different arrays without capturing again:RTX 5080, Julia 1.13. "Submit" is the CPU time to enqueue one iteration, "per iteration" is the throughput when submitting many iterations, and "latency" includes waiting for the result:
mul!+ 4 broadcasts)Eager execution is CPU-bound: the GPU sits idle between kernels, and the allocating RK4 code also suffers from GC pauses, hence its variable latency. With graphs, the GPU becomes the bottleneck: the CPU only spends half of each MLP iteration submitting work, and 5% of each RK4 iteration. The benchmark suite gains a
cuda/graphgroup tracking eager execution, graph launch and capture of 10 small kernels.Graphs lease the memory they use.
Managedmemory gets a lease count, and every way of releasing it (the pool, wrapped host memory,unsafe_free!) goes through a singlereleasegate that postpones the release until the last lease ends. While capturing,take_ownership!records the memory that captured operations use instead of synchronizing or changing its owning stream; the resultingCuGraph, and everyCuGraphExecinstantiated from it, each hold a lease. Launching an executable graph takes ownership of that memory like a kernel launch does, so other tasks synchronize correctly, and the eventual release is ordered after the last launch.Capturing no longer disables the GC. Finalizers only publish release requests (#3318). While a capture is in progress, retired resources aren't disposed of (doing so could add frees to the graph, or make prohibited API calls), and the capture drains them when it ends. Captures can't start while a disposal is in progress.
Captures don't get in the way of other tasks. Operations are captured on a dedicated non-blocking stream, which is the capturing task's stream for the duration of the capture, so other tasks can keep using the work the task submitted to its own stream earlier (e.g., an array it produced).
capture(f; stream)captures on a specific stream instead. Captures use the relaxed capture mode by default: the driver checks for unsafe API calls per thread rather than per task, so in the stricter modes an unrelated task that ran on the capturing thread while the capturing task was waiting (e.g., on a lock) had its allocations or synchronizations rejected, which invalidated the capture. CUDA.jl checks the operations it performs itself; the stricter modes remain available to debug libraries that don't support capture. Loading a kernel no longer synchronizes the device while a graph is being captured, and a capture that starts while loading a kernel synchronizes the device fails instead of waiting for it.Allocations during capture happen outside the graph. Graph-owned allocations are only valid between launches, which doesn't fit arrays whose lifetime the GC manages. Instead, memory is allocated from the regular pool on a separate stream, which owns that memory, so whoever uses it next (launching the graph, or another task) waits for the allocation like it would for any other operation. Scalars passed to libraries as
CuRefs are initialized right away rather than captured, so cuBLAS calls don't add memcpy nodes, which are expensive to launch alongside kernels.Host memory is snapshotted. A captured copy from CPU memory would read that memory on every launch, long after it may have been freed. Such copies are staged through pinned memory that the graph keeps alive, giving them snapshot semantics. Copying to CPU memory, waiting for the GPU, or using arrays from a
GPUArrays.@cachedscope throws a descriptiveCaptureErrorinstead of invalidating the capture. Arrays backed byHostMemorycan be used to exchange data with the CPU when launching a graph.Library allocations are handled. Libraries may allocate graph-owned memory during capture, which the driver doesn't free when the executable graph is updated or destroyed. Executable graphs track those allocations, free them when needed (like other memory, deferred while a capture is in progress), and are instantiated with
AUTO_FREE_ON_LAUNCHso they can be launched repeatedly.reclaim()also trims the graph memory pool.Limitations: random numbers generated within kernels (
rand()) or by the nativeCUDA.RNGare the same on every launch, because their seeds are kernel arguments;rand!/CUDA.rand, which use cuRAND, are not affected. Giving captured RNG kernels a device-side seed that a tiny kernel advances at the start of every replay works, but growsKernelState(changing every kernel's ABI) and doesn't cover the GPUArrays RNG, so that's better done separately. Seeding a cuRAND generator synchronizes the device (#3336), so a task that uses cuRAND for the first time, or re-seeds it, while another task is capturing invalidates that capture. Other tasks can't wait for work submitted to an explicitly passed capture stream before the capture until it ends. Tasks spawned during capture run their operations eagerly, outside the graph. Graphs captured directly through the driver API aren't integrated with the memory management.Tested on an RTX 5080 with Julia 1.12 (
core/graph,core/resource_graph).Fixes #3310, fixes #1230, fixes #2691. #1229 now results in a descriptive error.