Skip to content

Support waiting for completion notifications from the driver - #25

Merged
maleadt merged 3 commits into
mainfrom
tb/completion-notifications
Oct 1, 2026
Merged

maleadt merged 3 commits into
mainfrom
tb/completion-notifications

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

cooperative_wait (#23) waits for GPU operations without blocking the calling thread: it polls for a while, then hands a blocking driver wait to a worker thread. For GPUs that works well. Waiting takes as long as a plain blocking wait, and other tasks keep running.

It works badly for devices that execute on the host's own CPU cores, like PoCL (which is also KernelAbstractions' CPU() back-end) or Intel's CPU OpenCL runtime. KernelAbstractions' benchmark bot flagged this on JuliaGPU/KernelAbstractions.jl#813, with 15–30% slowdowns for saxpy kernels of 0.1–0.4 ms. I measured where the time goes with profiling markers on 4 cores:

PoCL, ~1 ms kernel (µs) total kernel execution kernel end → task resumes
blocking clWaitForEvents 1026–1041 1022–1035 2
cooperative_wait (worker) 1333–1402 1323–1392 5.6
completion callback (this PR) 1036–1039 1029–1032 3.5

The waking up afterwards isn't the problem. The kernel itself runs 25–35% longer. Per-thread scheduler statistics show that the kernel's threads get preempted three times as often: the worker is woken as the kernel starts, polling occupies a core, and the kernel then waits for its slowest thread. Intel's CPU runtime behaves the same. On an NVIDIA GPU the worker is as fast as blocking, but driver callbacks are useless there: NVIDIA's driver delivers a few percent of them about 20 ms late. So neither mechanism works for every back-end.

This PR makes cooperative_wait support both. A back-end can now pass subscribe(obj, payload), which registers a completion callback with the driver. The callback calls GPUToolbox.signal_completion(payload), and no worker thread is involved. With OpenCL, that looks like this:

function notify_completion(event::cl_event, status::Cint, payload::Ptr{Cvoid})
    GPUToolbox.signal_completion(payload)
    return
end
subscribe(event, payload) =
    clSetEventCallback(event, CL_COMPLETE,
                       @cfunction(notify_completion, Cvoid, (cl_event, Cint, Ptr{Cvoid})),
                       payload)

cooperative_wait(blocking_wait, event; subscribe, isdone=iscomplete, spin=10e-6)

OpenCL.jl (JuliaGPU/OpenCL.jl#518) uses this for CPU devices and keeps worker threads for GPUs. KernelAbstractions' POCL back-end uses it always, and CUDA.jl and oneAPI.jl keep using workers.

The part that took the most iterations is what the callback does on the driver's thread:

  • No detour through the event loop. It wakes the waiting task directly. My first version used uv_async_send and a dispatcher task, as the Julia manual suggests. That added a hop to another Julia thread, and that thread then spun in the scheduler into the next kernel, costing another ~35%.
  • Only a spin lock. The wakeup goes through Base.ThreadSynchronizer rather than a Base.Event, whose ReentrantLock could put the driver's thread to sleep in Julia's scheduler.
  • No finalizers on driver threads. Releasing Julia locks can run pending finalizers, which might call back into the driver while it holds a lock. They're deferred to a thread Julia manages. Debug builds of Julia are an exception; the docstring says so.
  • Registration can't be interrupted halfway. A waiter that gives up (a cancellable wait) parks its completion in a list until the driver has signalled it, so the driver never touches freed memory.

The docstring spells out the contract for back-ends: exactly one callback, possibly before subscribe returns, and GC-safe driver calls.

Polling also needed adjusting for these devices. The default polls for about 100 µs, first busy-waiting and then yielding, and that competes with the operation just like a worker does. Not polling at all makes tiny operations slower, though, because the callback has to wake the task. So spin now also accepts a duration (the third commit): busy-wait for at most that long. With spin=10e-6, KernelAbstractions' saxpy benchmark on 4 cores beats the old blocking wait for small kernels (3.1 vs 4.0 µs for 1024 elements, 12.2 vs 13.3 µs for 262144 elements). Longer operations are delayed by about half the polling time.

The first commit is a smaller, independent improvement: cooperative_wait now polls before allocating anything, so an operation found complete by polling costs no allocations on Julia before 1.14 (1.14's cancellation shield allocates a scope). The callback path allocates 160 bytes per wait, against 368 for a worker hand-off.

The new tests fire the callback from a thread that isn't managed by Julia, like a driver would. They cover:

  • callbacks during and after registration;
  • failing registration;
  • interrupts, with and without cancellable;
  • Julia 1.14 cancellation, again with and without cancellable;
  • allocation-free polling.

They pass on Julia 1.10, 1.12 and nightly. astra and sol reviewed several iterations of this. I also ran OpenCL.jl's synchronization, exception, array and memory tests on PoCL, Intel and NVIDIA, and KernelAbstractions' full suite, against this branch.

This bumps the version to 3.3.0, which JuliaGPU/OpenCL.jl#518 and JuliaGPU/KernelAbstractions.jl#813 depend on.

`cooperative_wait` allocated its wait state before polling, even though that
state is only needed when handing the wait to a worker. Poll first, keeping the
same guarantees: polling is shielded from cancellation, and for waits that
cannot be interrupted, an interrupt is only thrown once the operation has
completed, or once the worker's wait has, if polling gets interrupted.

Also specialize on the functions that are passed through, which Julia
otherwise doesn't, causing dynamic dispatch and boxing.
Waiting on a worker thread works well for GPUs, but not for devices that
execute on the host's CPU cores (e.g., PoCL, or Intel's CPU runtime): waking
the worker as the operation starts, and polling, compete with the operation for
those cores. Measured on 4 cores with PoCL, kernels of 0.2-1 ms took 25-35%
longer than with a plain blocking wait.

For those devices, `cooperative_wait` can now have the driver notify it when
the operation completes: `subscribe(obj, payload)` registers a driver callback,
which calls `GPUToolbox.signal_completion(payload)`. That directly wakes the
waiting task (going through the event loop instead would add a thread hop), so
no worker thread is involved, and kernels take as long as with a blocking wait.

Signalling only takes a spin lock, and defers finalizers to a thread managed by
Julia. Registration cannot be interrupted halfway, and a completion that its
waiter gave up on (e.g., because it was cancelled) is kept alive until the
driver has signalled it.
By default, `cooperative_wait` polls for a while, busy-waiting and then yielding
to other tasks, which takes around 100 µs. For devices that execute on the
host's CPU cores, that competes with the operation for those cores. Not polling
at all makes short operations slower, though: with KernelAbstractions' PoCL
back-end, trivial launches took 6.3 µs instead of 4.0 µs with a blocking wait.

So also accept a duration for `spin`, to only busy-wait for at most that long.
Polling for 10 µs before waiting for a completion notification, the saxpy
benchmark of KernelAbstractions on 4 cores takes 3.1 instead of 4.0 µs for 1024
elements, and 12.2 instead of 13.3 µs for 262144. Longer operations are delayed
by about half the polling time, which is why it should be kept short.
@maleadt
maleadt merged commit 3969374 into main Oct 1, 2026
10 checks passed
@maleadt
maleadt deleted the tb/completion-notifications branch October 1, 2026 05:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant