Support waiting for completion notifications from the driver - #25
Merged
Merged
Conversation
`cooperative_wait` allocated its wait state before polling, even though that state is only needed when handing the wait to a worker. Poll first, keeping the same guarantees: polling is shielded from cancellation, and for waits that cannot be interrupted, an interrupt is only thrown once the operation has completed, or once the worker's wait has, if polling gets interrupted. Also specialize on the functions that are passed through, which Julia otherwise doesn't, causing dynamic dispatch and boxing.
Waiting on a worker thread works well for GPUs, but not for devices that execute on the host's CPU cores (e.g., PoCL, or Intel's CPU runtime): waking the worker as the operation starts, and polling, compete with the operation for those cores. Measured on 4 cores with PoCL, kernels of 0.2-1 ms took 25-35% longer than with a plain blocking wait. For those devices, `cooperative_wait` can now have the driver notify it when the operation completes: `subscribe(obj, payload)` registers a driver callback, which calls `GPUToolbox.signal_completion(payload)`. That directly wakes the waiting task (going through the event loop instead would add a thread hop), so no worker thread is involved, and kernels take as long as with a blocking wait. Signalling only takes a spin lock, and defers finalizers to a thread managed by Julia. Registration cannot be interrupted halfway, and a completion that its waiter gave up on (e.g., because it was cancelled) is kept alive until the driver has signalled it.
By default, `cooperative_wait` polls for a while, busy-waiting and then yielding to other tasks, which takes around 100 µs. For devices that execute on the host's CPU cores, that competes with the operation for those cores. Not polling at all makes short operations slower, though: with KernelAbstractions' PoCL back-end, trivial launches took 6.3 µs instead of 4.0 µs with a blocking wait. So also accept a duration for `spin`, to only busy-wait for at most that long. Polling for 10 µs before waiting for a completion notification, the saxpy benchmark of KernelAbstractions on 4 cores takes 3.1 instead of 4.0 µs for 1024 elements, and 12.2 instead of 13.3 µs for 262144. Longer operations are delayed by about half the polling time, which is why it should be kept short.
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
cooperative_wait(#23) waits for GPU operations without blocking the calling thread: it polls for a while, then hands a blocking driver wait to a worker thread. For GPUs that works well. Waiting takes as long as a plain blocking wait, and other tasks keep running.It works badly for devices that execute on the host's own CPU cores, like PoCL (which is also KernelAbstractions'
CPU()back-end) or Intel's CPU OpenCL runtime. KernelAbstractions' benchmark bot flagged this on JuliaGPU/KernelAbstractions.jl#813, with 15–30% slowdowns for saxpy kernels of 0.1–0.4 ms. I measured where the time goes with profiling markers on 4 cores:clWaitForEventscooperative_wait(worker)The waking up afterwards isn't the problem. The kernel itself runs 25–35% longer. Per-thread scheduler statistics show that the kernel's threads get preempted three times as often: the worker is woken as the kernel starts, polling occupies a core, and the kernel then waits for its slowest thread. Intel's CPU runtime behaves the same. On an NVIDIA GPU the worker is as fast as blocking, but driver callbacks are useless there: NVIDIA's driver delivers a few percent of them about 20 ms late. So neither mechanism works for every back-end.
This PR makes
cooperative_waitsupport both. A back-end can now passsubscribe(obj, payload), which registers a completion callback with the driver. The callback callsGPUToolbox.signal_completion(payload), and no worker thread is involved. With OpenCL, that looks like this:OpenCL.jl (JuliaGPU/OpenCL.jl#518) uses this for CPU devices and keeps worker threads for GPUs. KernelAbstractions' POCL back-end uses it always, and CUDA.jl and oneAPI.jl keep using workers.
The part that took the most iterations is what the callback does on the driver's thread:
uv_async_sendand a dispatcher task, as the Julia manual suggests. That added a hop to another Julia thread, and that thread then spun in the scheduler into the next kernel, costing another ~35%.Base.ThreadSynchronizerrather than aBase.Event, whoseReentrantLockcould put the driver's thread to sleep in Julia's scheduler.The docstring spells out the contract for back-ends: exactly one callback, possibly before
subscribereturns, and GC-safe driver calls.Polling also needed adjusting for these devices. The default polls for about 100 µs, first busy-waiting and then yielding, and that competes with the operation just like a worker does. Not polling at all makes tiny operations slower, though, because the callback has to wake the task. So
spinnow also accepts a duration (the third commit): busy-wait for at most that long. Withspin=10e-6, KernelAbstractions' saxpy benchmark on 4 cores beats the old blocking wait for small kernels (3.1 vs 4.0 µs for 1024 elements, 12.2 vs 13.3 µs for 262144 elements). Longer operations are delayed by about half the polling time.The first commit is a smaller, independent improvement:
cooperative_waitnow polls before allocating anything, so an operation found complete by polling costs no allocations on Julia before 1.14 (1.14's cancellation shield allocates a scope). The callback path allocates 160 bytes per wait, against 368 for a worker hand-off.The new tests fire the callback from a thread that isn't managed by Julia, like a driver would. They cover:
cancellable;cancellable;They pass on Julia 1.10, 1.12 and nightly. astra and sol reviewed several iterations of this. I also ran OpenCL.jl's synchronization, exception, array and memory tests on PoCL, Intel and NVIDIA, and KernelAbstractions' full suite, against this branch.
This bumps the version to 3.3.0, which JuliaGPU/OpenCL.jl#518 and JuliaGPU/KernelAbstractions.jl#813 depend on.