Skip to content

Use GPUToolbox's cooperative_wait for synchronization - #1118

Open
maleadt wants to merge 5 commits into
mainfrom
tb/cooperative-wait
Open

maleadt wants to merge 5 commits into
mainfrom
tb/cooperative-wait

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

CUDA.jl, oneAPI.jl and OpenCL.jl wait for the GPU on a worker thread, so that the calling task doesn't block its thread. That implementation now lives in GPUToolbox as cooperative_wait (JuliaGPU/GPUToolbox.jl#23). This PR switches AMDGPU.jl over to it for streams, events and HIP.device_synchronize().

Until now, nonblocking stream synchronization enqueued a host function with hipLaunchHostFunc that woke the waiting task through libuv. Because that host function might never run if something went wrong, every synchronization also spawned a second task that polled the stream on a 1 s timer. cooperative_wait polls the stream first, exactly like before, and then hands a plain hipStreamSynchronize to a small pool of worker threads while the calling task waits on an event. Errors come back from hipStreamSynchronize itself, so the watchdog goes away. This also fixes a few other things:

  • Host functions are slow and get in the way of the stream. CLR puts a blocking marker behind the callback, so later work on the stream waits for the host function too, and the wakeup then goes through the libuv event loop. On CUDA, the equivalent cuLaunchHostFunc path added 20–75 µs over a blocking wait, against 2–4 µs for handing the wait to a worker (measured with CUDA.jl on an RTX 5080).

  • Event synchronization blocked the thread. synchronize(::HIPEvent) polled for a while and then called hipEventSynchronize on the calling thread. It now waits on a worker too.

  • HIP.device_synchronize() blocked the thread, and could deadlock with hostcalls. The host side of a hostcall is a task. If a kernel waits for a hostcall whose task can only run on the thread blocked in hipDeviceSynchronize, nothing makes progress. For example, with julia --threads=1:

    hc = AMDGPU.Device.HostCallHolder(Float32, Tuple{Float32}; continuous=true) do x
        x + 42f0
    end
    kernel!(y, hc) = (y[1] = AMDGPU.Device.hostcall!(hc, y[1]); nothing)
    y = ROCArray(Float32[0])
    @roc kernel!(y, hc)
    AMDGPU.HIP.device_synchronize()   # hangs on main, returns with this PR
  • Every synchronization called device! from the watchdog task, which also sets the global default device.

The workers call HIP directly rather than through the checked wrappers, because those switch to the context of the task making the call. They select the caller's device with hipSetDevice first, since the null stream and hipDeviceSynchronize depend on it. blocking=true and the nonblocking_synchronization preference work as before. Finalizers, which can't switch tasks, keep the context they selected and wait on the calling thread.

This requires GPUToolbox 3.3.1, which fixes two ways the workers could deadlock (JuliaGPU/GPUToolbox.jl#26). Both matter for the hostcall case above. hipDeviceSynchronize can't be polled, and before 3.3.1 such a wait had to wait for a free worker when all of them were busy, which never happens if the busy workers are waiting for kernels that need the hostcall to be served. And a worker could run a finalizer that blocks (e.g. freeing memory, which can wait for the device to go idle) before waking up the task it waited for. Now such waits get an extra worker, and workers never run finalizers.

This hasn't been run on a discrete AMD GPU yet, so CI is the real validation. I did compare both implementations on the integrated GPU of a Ryzen 9 9950X (gfx1036, run as gfx1030 with ROCm 7.2.4), in one process, alternating between them in blocks. The tables show the median (p90) time for a launch plus synchronize() of a busy-waiting kernel:

--threads=1:

kernel blocking before after
0 µs 6.7 (7.3) µs 8.8 (12.1) µs 8.7 (11.6) µs
100 µs 112.2 (112.9) µs 117.4 (120.4) µs 113.8 (116.1) µs
1 ms 1018.5 (1022.5) µs 1027.4 (1057.9) µs 1021.2 (1037.9) µs
5 ms 5026.2 (5050.9) µs 5068.9 (5119.9) µs 5050.7 (5083.0) µs

--threads=4:

kernel blocking before after
0 µs 6.8 (7.2) µs 7.0 (8.3) µs 7.0 (7.8) µs
100 µs 111.2 (111.6) µs 122.2 (128.6) µs 112.9 (114.6) µs
1 ms 1017.3 (1018.2) µs 1044.7 (1068.0) µs 1020.3 (1035.1) µs
5 ms 5025.0 (5047.6) µs 5084.9 (5169.0) µs 5044.6 (5083.1) µs

Short kernels finish while polling, so they aren't affected. For longer ones, the new version was never slower than the host-function path in any run, and usually had a much tighter tail. The machine was heavily loaded, so the 5 ms numbers varied between runs, but the ordering didn't change.

The trade-off is CPU usage while waiting. HIP busy-waits in hipStreamSynchronize, so the worker keeps a core busy, just like blocking=true does on the calling thread. On this machine, polling a stream with hipStreamQuery also kept one of HSA's runtime threads busy, both before and after this PR. As a result, a long wait used about 2 cores instead of about 1. HIP.synchronize(stream; spin=false) skips the polling and brings that back to 1 core, at the cost of a few µs on short kernels.

The new tests check that synchronize(stream) (including the null stream and spin=false), synchronize(event) and HIP.device_synchronize() don't block the thread, and that a hostcall gets served while device_synchronize() waits in a single-threaded process. They don't measure anything. A kernel spins until a flag in host memory is set, and another task on the same thread sets it after yielding many more times than the polling before a wait does, so the wait can only return if that task got to run in the meantime. If it never does, the kernel gives up after 20 s or so and records that it timed out, so a regression fails instead of hanging. That also makes the tests indifferent to the MI250 runner stalling test processes (individual tests there regularly take 5–30× longer than usual, also on main). On main, the event and device tests fail, and the hostcall test times out. With GPUToolbox 3.3.2, they pass on the integrated GPU mentioned above, also with --threads=1 and while the process is repeatedly stopped and resumed, and with nonblocking_synchronization=false they fail as expected.

Nonblocking stream synchronization enqueued a host function that woke the
waiting task through libuv, plus a task polling the stream every second in
case the host function never ran. Instead, wait for the stream in HIP on a
worker thread, as CUDA.jl, oneAPI.jl and OpenCL.jl do with the shared
implementation in GPUToolbox. This avoids the host function, which delays
both the wakeup and later work on the stream, and the extra tasks and
timer per synchronization.

Event synchronization used to block the calling thread after polling, and
now waits the same way.
HIP.device_synchronize called hipDeviceSynchronize on the calling thread,
which deadlocked when a kernel was waiting for a hostcall whose task could
only run on that thread. Wait on a worker thread instead, like stream and
event synchronization do.
HIP.device_synchronize() waits for something that cannot be polled. With
earlier versions, cooperative_wait made such waits wait for a free worker
when all of them were busy, and workers could get stuck running finalizers
before waking up the waiting task. Both could deadlock when the GPU work
depends on the waiting task's thread making progress, as with hostcalls.
The tests launched a kernel and then checked that other tasks made progress
while waiting for it. On a loaded node (the MI250 runner regularly runs tests
5-30x slower than usual, also on main), the process can be descheduled for
longer than the kernel takes, e.g. while compiling the measurement after the
launch. The kernel then completed before the wait started, which looked like
a blocked thread. The wall-clock calibration was affected too, picking
kernels that took only a few ms.

Time the calibration kernel on the GPU, compile everything before launching
the kernel, and only accept a measurement when the wait took a while,
retrying with a longer kernel otherwise.
The previous tests counted how often another task ran while waiting for a
kernel, which depends on timing: when the process gets descheduled for longer
than the kernel takes, as happens on the MI250 runner, the kernel completes
before the wait starts, which looks like a blocked thread. Retrying with longer
kernels made that less likely, but not impossible.

Instead, keep the GPU busy with a kernel that spins until a flag in host memory
is set, and set that flag from another task on the same thread, after it has
yielded many more times than the polling at the start of a wait does. The wait
can then only return if that task ran in the meantime. If it doesn't, the
kernel gives up after a while and records that it timed out, so a blocking
implementation fails instead of hanging. `blocking=true` is checked the same
way, expecting the kernel to time out.

@gbaraldi gbaraldi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested on MI300A + MI250 (ROCm 7.2.4), 1 and 4 threads. The new tests pass (on main the event/device tests fail and the hostcall test hangs), multi-GPU and a 64-task stress run are OK, and latency is ≤ main with tighter tails.

  1. Overflow threads. device_synchronize has no isdone, so each call made while the workers are busy gets a new overflow thread, and those threads are never torn down. 64 concurrent calls left 42 extra spinning threads. Allow one sync per device at a time, or poll an event on the null stream.
  2. Ctrl-C during synchronize waits until the kernel finishes (12 s vs 0.5 s on main), so a hung kernel can't be interrupted. Document it or expose cancellable.
  3. CPU use. Once any host function has been used in the process (e.g. #1116), polling with hipStreamQuery makes a CLR thread spin: 2 cores per wait instead of 1 (spin=false: 1).
  4. Nits: spin=false costs 15–20 µs on short kernels; spin isn't documented or passed through AMDGPU.synchronize; the device_synchronize docstring still says "Blocks"; docs/src/api/system.md says nonblocking_synchronize (predates the PR).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants