Wait for the device cooperatively - #518
Merged
Merged
Conversation
maleadt
force-pushed
the
tb/cooperative-wait
branch
from
September 30, 2026 18:59
f32a982 to
a68f55d
Compare
Many OpenCL calls can block for a long time: waiting for commands, compiling programs, blocking transfers. As plain ccalls, they prevent every other thread from collecting garbage in the meantime. They can even deadlock: if a driver thread enters Julia (e.g., to run a callback) while a Julia thread blocks in the driver waiting on it, the GC waits on the blocked thread, and the driver thread waits on the GC (see JuliaGPU/CUDA.jl#2261). As in CUDA.jl, make all API calls with GPUToolbox's `@gcsafe_ccall`, including extension functions, which are called through function pointers (supported since GPUToolbox 3.2). On Julia versions without native GC-safe calls, `@gcsafe_ccall` converts arguments itself before passing them to `ccall`, which converts them again. Make that second, identity conversion unambiguous for `PtrOrCLPtr` and `RefOrCLRef`, as CUDA.jl does.
`wait(::cl.AbstractEvent)` called `clWaitForEvents`, blocking the calling thread until the commands completed. No other task could run on that thread in the meantime, and waiting could not be interrupted. Use GPUToolbox's `cooperative_wait` instead, like CUDA.jl does. It polls the execution status of the events for a while, which keeps short waits fast, and then hands a blocking `clWaitForEvents` to a worker thread while the task yields. Since polling only makes progress once commands have been submitted, the events' queues are flushed first. Afterwards, `clWaitForEvents` is called again (returning immediately) to synchronize host memory and to report errors. Waiting falls back to blocking in finalizers and while precompiling, and the `nonblocking_synchronization` preference switches back to blocking entirely.
`cl.finish`, and with it `OpenCL.synchronize`, `OpenCL.@sync` and `KernelAbstractions.synchronize`, blocked the calling thread in `clFinish`. KernelAbstractions requires `synchronize` to yield instead, e.g., to overlap MPI communication with computation, and for `KA.@spawn` tasks to run concurrently. OpenCL cannot query whether a queue is idle, so wait cooperatively for a marker instead, which completes once all previously submitted commands have. The trailing `clFinish` is gone too: other tasks may submit more work to the queue while we wait, and this is what `synchronize` promises anyway. When checking for device-side exceptions, wait for the queue before taking the mailbox lock, so that other tasks can keep launching kernels in the meantime. Under the lock, the pending queues still need to be waited for before the mailbox can be read. That includes this queue, but only if a kernel was launched since we waited for it, which the launch counter (now atomic, and only published once a launch has been submitted) tells us. Mailboxes that need to be mapped for host access now do so on a queue of their own: mapping on the synchronized queue would also wait, blocking the thread, for any work other tasks submitted to it in the meantime.
Blocking reads, writes and copies (e.g., `Array(x)`, or `copyto!` to host memory) passed `blocking=true` to the driver, which then blocked the calling thread until all earlier work on the queue had completed too. Submit these commands without blocking in the driver, and wait for their event cooperatively instead. That wait cannot be interrupted: the command accesses memory that the caller may release as soon as we return. This also makes `enqueue_copy` between buffers honor `blocking`, which it ignored before. Maps still block in the driver, as they are only issued after synchronizing.
maleadt
force-pushed
the
tb/cooperative-wait
branch
from
September 30, 2026 19:25
a68f55d to
b385ef7
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #518 +/- ##
==========================================
+ Coverage 85.53% 85.56% +0.02%
==========================================
Files 19 19
Lines 1680 1683 +3
==========================================
+ Hits 1437 1440 +3
Misses 243 243 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Handing the blocking wait to a worker thread works well for GPUs, but slows down devices that execute on the host's CPU cores (e.g., PoCL, or Intel's CPU runtime): waking the worker as the commands start, and polling their status for a long time, compete with them for those cores. On 4 cores with PoCL, kernels of 0.2-1 ms took 25-35% longer than with a blocking wait. For such devices, only poll for at most 10 µs, and then have the driver notify us when the commands complete, using `clSetEventCallback`, which GPUToolbox 3.3 supports. Kernels then take about as long as with a blocking wait. GPUs keep using worker threads: NVIDIA's driver delivers a few percent of these callbacks about 20 ms late.
On GPUs, waits are handed to GPUToolbox's worker threads, which 3.3.1 keeps from deadlocking: workers no longer run finalizers, which could block (e.g., freeing memory waits for the device) before the worker notifies its waiter (JuliaGPU/GPUToolbox.jl#26).
On CPU devices, which wait for completion callbacks, 3.3.2 keeps the woken task from spinning on a lock held by the driver's callback thread after preempting it, which made longer PoCL kernels slower when all cores are busy (JuliaGPU/GPUToolbox.jl#27).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Whenever OpenCL.jl waited for the device, it blocked the calling thread inside the OpenCL driver. That covers
OpenCL.synchronize,KernelAbstractions.synchronize,waiton an event, and blocking copies likeArray(x), which all ended up inclFinishorclWaitForEvents. Other tasks scheduled on that thread couldn't run until the device was done, and ^C did nothing.That's a problem beyond responsiveness: KernelAbstractions explicitly requires
synchronizeto be cooperative. MPI codes rely on it to overlap communication with computation, and without it, tasks started withKA.@spawndon't actually run concurrently.With this PR, all of those waits yield to the Julia scheduler. Take a kernel that runs for 600–850 ms, and a task on the same thread that ticks every 10 ms:
OpenCL.synchronize()Array(a)OpenCL.synchronize()Array(a)What changes for users
Mostly nothing, other than the waiting task no longer hogging its thread. A few things are worth knowing:
synchronizeorwait(event)now cancels the wait and returns to the prompt. The device keeps executing whatever it was running.Array(x),copyto!to host memory, ...) also yield while waiting, but can't be interrupted before the copy finishes. The caller may free or reuse the host memory as soon as the copy returns, so it must not return early.cl.enqueue_copy(...; blocking=true)between two buffers used to ignoreblocking, and now honors it.nonblocking_synchronizationpreference tofalserestores the old blocking behavior, e.g. for bisecting.How it works
The waiting is done by
GPUToolbox.cooperative_wait, which CUDA.jl uses too, so the back-ends share one implementation. How it waits depends on the device:clWaitForEventsto a small pool of worker threads, and let the waiting task sleep until the worker is done.clSetEventCallback). Here the device runs on the host's own cores, and the GPU approach made kernels 25–35% slower: waking a worker thread just as the kernel starts, and polling for long, both take cores away from the kernel. Callbacks involve no extra thread, and kernels run about as fast as with a blocking wait.Why not callbacks everywhere? NVIDIA's driver delivers a few percent of them about 20 ms late. GPUToolbox gained support for callbacks for this in JuliaGPU/GPUToolbox.jl#25.
Either way, OpenCL.jl afterwards calls
clWaitForEventsonce more. That returns immediately at that point, but it still synchronizes host memory and reports failed commands exactly as before.OpenCL has no call to ask whether a queue is idle, so
synchronizeenqueues a marker (a no-op command that completes once everything before it has) and waits for that. The device-exception check, which runs as part of everysynchronize, now waits before taking its lock. Other tasks can therefore keep launching kernels while one task is waiting. It still makes sure no kernel can write to the exception mailbox while it reads it.This PR also makes every OpenCL API call GC-safe, as CUDA.jl did in JuliaGPU/CUDA.jl#2262. Otherwise, a thread blocked in the driver, for example compiling a program or waiting for a command, keeps all other threads from collecting garbage. The worker threads need this too, since they sit in
clWaitForEventsfor as long as a command runs. That includes extension functions, which are called through function pointers (supported by@gcsafe_ccallsince JuliaGPU/GPUToolbox.jl#24).Performance
Waiting cooperatively costs little. Launching kernels of different durations and waiting for them, median over 200–300 runs, on 4 cores:
On NVIDIA, cooperative waits are within a few µs of blocking ones at all sizes.
The one thing that got slower is
synchronizeon a queue that is already idle: it now costs about 5 µs on NVIDIA and 13 µs on PoCL, whereclFinishtook less than a microsecond, because of the marker.Testing
The new tests make the device wait for a user event that another task sets a little later. With a thread-blocking wait, that task never gets to run, so the tests would deadlock. They cover waiting for single events and lists of events,
synchronize, blocking copies, and failed commands. On Julia 1.14 they also check that cancellation abandons an event wait but not a blocking copy.The synchronization, exceptions, array, memory, buffer, execution and KernelAbstractions tests pass on PoCL, Intel's CPU runtime and NVIDIA. Intel also crashes in two GPUArrays tests (
linalg/coreandreductions/mapreducedim!), butmaindoes the same.This needs GPUToolbox 3.3.2. 3.3 added the completion callbacks (JuliaGPU/GPUToolbox.jl#25), 3.3.1 fixes two ways in which the worker threads we use for GPUs could deadlock (JuliaGPU/GPUToolbox.jl#26), and 3.3.2 fixes a problem with the callbacks used for CPU devices: the woken task could spin on a lock held by the driver's callback thread after preempting it, which made longer PoCL kernels slower when all cores are busy (JuliaGPU/GPUToolbox.jl#27). KernelAbstractions' own PoCL back-end gets the same change in JuliaGPU/KernelAbstractions.jl#813.