Skip to content

Wait for the device cooperatively - #518

Merged
maleadt merged 7 commits into
mainfrom
tb/cooperative-wait
Oct 2, 2026
Merged

maleadt merged 7 commits into
mainfrom
tb/cooperative-wait

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Whenever OpenCL.jl waited for the device, it blocked the calling thread inside the OpenCL driver. That covers OpenCL.synchronize, KernelAbstractions.synchronize, wait on an event, and blocking copies like Array(x), which all ended up in clFinish or clWaitForEvents. Other tasks scheduled on that thread couldn't run until the device was done, and ^C did nothing.

That's a problem beyond responsiveness: KernelAbstractions explicitly requires synchronize to be cooperative. MPI codes rely on it to overlap communication with computation, and without it, tasks started with KA.@spawn don't actually run concurrently.

With this PR, all of those waits yield to the Julia scheduler. Take a kernel that runs for 600–850 ms, and a task on the same thread that ticks every 10 ms:

ticker = @async while !done[]
    ticks[] += 1
    sleep(0.01)
end
@opencl global_size=length(a) slow_kernel(a)
OpenCL.synchronize()  # or Array(a), or wait(event)
before after
PoCL, OpenCL.synchronize() 0 ticks 75 ticks
PoCL, Array(a) 0 ticks 32 ticks
NVIDIA, OpenCL.synchronize() 0 ticks 58 ticks
NVIDIA, Array(a) 0 ticks 39 ticks

What changes for users

Mostly nothing, other than the waiting task no longer hogging its thread. A few things are worth knowing:

  • On Julia 1.14, ^C during synchronize or wait(event) now cancels the wait and returns to the prompt. The device keeps executing whatever it was running.
  • Blocking copies (Array(x), copyto! to host memory, ...) also yield while waiting, but can't be interrupted before the copy finishes. The caller may free or reuse the host memory as soon as the copy returns, so it must not return early.
  • Finalizers and precompilation still block, since neither can yield.
  • cl.enqueue_copy(...; blocking=true) between two buffers used to ignore blocking, and now honors it.
  • Setting the nonblocking_synchronization preference to false restores the old blocking behavior, e.g. for bisecting.

How it works

The waiting is done by GPUToolbox.cooperative_wait, which CUDA.jl uses too, so the back-ends share one implementation. How it waits depends on the device:

  • GPUs: poll the command's execution status for a short while, which keeps short waits as fast as before. If the command is still running after that, hand a blocking clWaitForEvents to a small pool of worker threads, and let the waiting task sleep until the worker is done.
  • CPU devices (PoCL, Intel's CPU runtime): poll for at most 10 µs, then have the driver call us back when the command completes (clSetEventCallback). Here the device runs on the host's own cores, and the GPU approach made kernels 25–35% slower: waking a worker thread just as the kernel starts, and polling for long, both take cores away from the kernel. Callbacks involve no extra thread, and kernels run about as fast as with a blocking wait.

Why not callbacks everywhere? NVIDIA's driver delivers a few percent of them about 20 ms late. GPUToolbox gained support for callbacks for this in JuliaGPU/GPUToolbox.jl#25.

Either way, OpenCL.jl afterwards calls clWaitForEvents once more. That returns immediately at that point, but it still synchronizes host memory and reports failed commands exactly as before.

OpenCL has no call to ask whether a queue is idle, so synchronize enqueues a marker (a no-op command that completes once everything before it has) and waits for that. The device-exception check, which runs as part of every synchronize, now waits before taking its lock. Other tasks can therefore keep launching kernels while one task is waiting. It still makes sure no kernel can write to the exception mailbox while it reads it.

This PR also makes every OpenCL API call GC-safe, as CUDA.jl did in JuliaGPU/CUDA.jl#2262. Otherwise, a thread blocked in the driver, for example compiling a program or waiting for a command, keeps all other threads from collecting garbage. The worker threads need this too, since they sit in clWaitForEvents for as long as a command runs. That includes extension functions, which are called through function pointers (supported by @gcsafe_ccall since JuliaGPU/GPUToolbox.jl#24).

Performance

Waiting cooperatively costs little. Launching kernels of different durations and waiting for them, median over 200–300 runs, on 4 cores:

PoCL, total µs ~0 µs kernel ~20 µs ~50 µs ~200 µs ~1 ms
blocking 12–13 14 57 226–228 1030–1035
cooperative 13 16 63 237–238 1050–1053

On NVIDIA, cooperative waits are within a few µs of blocking ones at all sizes.

The one thing that got slower is synchronize on a queue that is already idle: it now costs about 5 µs on NVIDIA and 13 µs on PoCL, where clFinish took less than a microsecond, because of the marker.

Testing

The new tests make the device wait for a user event that another task sets a little later. With a thread-blocking wait, that task never gets to run, so the tests would deadlock. They cover waiting for single events and lists of events, synchronize, blocking copies, and failed commands. On Julia 1.14 they also check that cancellation abandons an event wait but not a blocking copy.

The synchronization, exceptions, array, memory, buffer, execution and KernelAbstractions tests pass on PoCL, Intel's CPU runtime and NVIDIA. Intel also crashes in two GPUArrays tests (linalg/core and reductions/mapreducedim!), but main does the same.

This needs GPUToolbox 3.3.2. 3.3 added the completion callbacks (JuliaGPU/GPUToolbox.jl#25), 3.3.1 fixes two ways in which the worker threads we use for GPUs could deadlock (JuliaGPU/GPUToolbox.jl#26), and 3.3.2 fixes a problem with the callbacks used for CPU devices: the woken task could spin on a lock held by the driver's callback thread after preempting it, which made longer PoCL kernels slower when all cores are busy (JuliaGPU/GPUToolbox.jl#27). KernelAbstractions' own PoCL back-end gets the same change in JuliaGPU/KernelAbstractions.jl#813.

Many OpenCL calls can block for a long time: waiting for commands, compiling
programs, blocking transfers. As plain ccalls, they prevent every other thread
from collecting garbage in the meantime. They can even deadlock: if a driver
thread enters Julia (e.g., to run a callback) while a Julia thread blocks in
the driver waiting on it, the GC waits on the blocked thread, and the driver
thread waits on the GC (see JuliaGPU/CUDA.jl#2261).

As in CUDA.jl, make all API calls with GPUToolbox's `@gcsafe_ccall`, including
extension functions, which are called through function pointers (supported
since GPUToolbox 3.2).

On Julia versions without native GC-safe calls, `@gcsafe_ccall` converts
arguments itself before passing them to `ccall`, which converts them again.
Make that second, identity conversion unambiguous for `PtrOrCLPtr` and
`RefOrCLRef`, as CUDA.jl does.
`wait(::cl.AbstractEvent)` called `clWaitForEvents`, blocking the calling
thread until the commands completed. No other task could run on that thread
in the meantime, and waiting could not be interrupted.

Use GPUToolbox's `cooperative_wait` instead, like CUDA.jl does. It polls the
execution status of the events for a while, which keeps short waits fast, and
then hands a blocking `clWaitForEvents` to a worker thread while the task
yields. Since polling only makes progress once commands have been submitted,
the events' queues are flushed first. Afterwards, `clWaitForEvents` is called
again (returning immediately) to synchronize host memory and to report errors.

Waiting falls back to blocking in finalizers and while precompiling, and the
`nonblocking_synchronization` preference switches back to blocking entirely.
`cl.finish`, and with it `OpenCL.synchronize`, `OpenCL.@sync` and
`KernelAbstractions.synchronize`, blocked the calling thread in `clFinish`.
KernelAbstractions requires `synchronize` to yield instead, e.g., to overlap
MPI communication with computation, and for `KA.@spawn` tasks to run
concurrently.

OpenCL cannot query whether a queue is idle, so wait cooperatively for a
marker instead, which completes once all previously submitted commands have.
The trailing `clFinish` is gone too: other tasks may submit more work to the
queue while we wait, and this is what `synchronize` promises anyway.

When checking for device-side exceptions, wait for the queue before taking the
mailbox lock, so that other tasks can keep launching kernels in the meantime.
Under the lock, the pending queues still need to be waited for before the
mailbox can be read. That includes this queue, but only if a kernel was
launched since we waited for it, which the launch counter (now atomic, and
only published once a launch has been submitted) tells us.

Mailboxes that need to be mapped for host access now do so on a queue of their
own: mapping on the synchronized queue would also wait, blocking the thread,
for any work other tasks submitted to it in the meantime.
Blocking reads, writes and copies (e.g., `Array(x)`, or `copyto!` to host
memory) passed `blocking=true` to the driver, which then blocked the calling
thread until all earlier work on the queue had completed too.

Submit these commands without blocking in the driver, and wait for their event
cooperatively instead. That wait cannot be interrupted: the command accesses
memory that the caller may release as soon as we return. This also makes
`enqueue_copy` between buffers honor `blocking`, which it ignored before.
Maps still block in the driver, as they are only issued after synchronizing.
@maleadt
maleadt force-pushed the tb/cooperative-wait branch from a68f55d to b385ef7 Compare September 30, 2026 19:25
@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 85.56%. Comparing base (e3f48f3) to head (3f57841).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #518      +/-   ##
==========================================
+ Coverage   85.53%   85.56%   +0.02%     
==========================================
  Files          19       19              
  Lines        1680     1683       +3     
==========================================
+ Hits         1437     1440       +3     
  Misses        243      243              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Handing the blocking wait to a worker thread works well for GPUs, but slows
down devices that execute on the host's CPU cores (e.g., PoCL, or Intel's CPU
runtime): waking the worker as the commands start, and polling their status for
a long time, compete with them for those cores. On 4 cores with PoCL, kernels of
0.2-1 ms took 25-35% longer than with a blocking wait.

For such devices, only poll for at most 10 µs, and then have the driver notify
us when the commands complete, using `clSetEventCallback`, which GPUToolbox 3.3
supports. Kernels then take about as long as with a blocking wait. GPUs keep
using worker threads: NVIDIA's driver delivers a few percent of these callbacks
about 20 ms late.
On GPUs, waits are handed to GPUToolbox's worker threads, which 3.3.1 keeps
from deadlocking: workers no longer run finalizers, which could block (e.g.,
freeing memory waits for the device) before the worker notifies its waiter
(JuliaGPU/GPUToolbox.jl#26).
On CPU devices, which wait for completion callbacks, 3.3.2 keeps the woken
task from spinning on a lock held by the driver's callback thread after
preempting it, which made longer PoCL kernels slower when all cores are busy
(JuliaGPU/GPUToolbox.jl#27).
@maleadt
maleadt merged commit 6362ee0 into main Oct 2, 2026
16 checks passed
@maleadt
maleadt deleted the tb/cooperative-wait branch October 2, 2026 05:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant