Conversation
Nonblocking stream synchronization enqueued a host function that woke the waiting task through libuv, plus a task polling the stream every second in case the host function never ran. Instead, wait for the stream in HIP on a worker thread, as CUDA.jl, oneAPI.jl and OpenCL.jl do with the shared implementation in GPUToolbox. This avoids the host function, which delays both the wakeup and later work on the stream, and the extra tasks and timer per synchronization. Event synchronization used to block the calling thread after polling, and now waits the same way.
HIP.device_synchronize called hipDeviceSynchronize on the calling thread, which deadlocked when a kernel was waiting for a hostcall whose task could only run on that thread. Wait on a worker thread instead, like stream and event synchronization do.
HIP.device_synchronize() waits for something that cannot be polled. With earlier versions, cooperative_wait made such waits wait for a free worker when all of them were busy, and workers could get stuck running finalizers before waking up the waiting task. Both could deadlock when the GPU work depends on the waiting task's thread making progress, as with hostcalls.
The tests launched a kernel and then checked that other tasks made progress while waiting for it. On a loaded node (the MI250 runner regularly runs tests 5-30x slower than usual, also on main), the process can be descheduled for longer than the kernel takes, e.g. while compiling the measurement after the launch. The kernel then completed before the wait started, which looked like a blocked thread. The wall-clock calibration was affected too, picking kernels that took only a few ms. Time the calibration kernel on the GPU, compile everything before launching the kernel, and only accept a measurement when the wait took a while, retrying with a longer kernel otherwise.
The previous tests counted how often another task ran while waiting for a kernel, which depends on timing: when the process gets descheduled for longer than the kernel takes, as happens on the MI250 runner, the kernel completes before the wait starts, which looks like a blocked thread. Retrying with longer kernels made that less likely, but not impossible. Instead, keep the GPU busy with a kernel that spins until a flag in host memory is set, and set that flag from another task on the same thread, after it has yielded many more times than the polling at the start of a wait does. The wait can then only return if that task ran in the meantime. If it doesn't, the kernel gives up after a while and records that it timed out, so a blocking implementation fails instead of hanging. `blocking=true` is checked the same way, expecting the kernel to time out.
gbaraldi
reviewed
Oct 1, 2026
gbaraldi
left a comment
Member
There was a problem hiding this comment.
Tested on MI300A + MI250 (ROCm 7.2.4), 1 and 4 threads. The new tests pass (on main the event/device tests fail and the hostcall test hangs), multi-GPU and a 64-task stress run are OK, and latency is ≤ main with tighter tails.
- Overflow threads.
device_synchronizehas noisdone, so each call made while the workers are busy gets a new overflow thread, and those threads are never torn down. 64 concurrent calls left 42 extra spinning threads. Allow one sync per device at a time, or poll an event on the null stream. - Ctrl-C during
synchronizewaits until the kernel finishes (12 s vs 0.5 s on main), so a hung kernel can't be interrupted. Document it or exposecancellable. - CPU use. Once any host function has been used in the process (e.g. #1116), polling with
hipStreamQuerymakes a CLR thread spin: 2 cores per wait instead of 1 (spin=false: 1). - Nits:
spin=falsecosts 15–20 µs on short kernels;spinisn't documented or passed throughAMDGPU.synchronize; thedevice_synchronizedocstring still says "Blocks";docs/src/api/system.mdsaysnonblocking_synchronize(predates the PR).
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
CUDA.jl, oneAPI.jl and OpenCL.jl wait for the GPU on a worker thread, so that the calling task doesn't block its thread. That implementation now lives in GPUToolbox as
cooperative_wait(JuliaGPU/GPUToolbox.jl#23). This PR switches AMDGPU.jl over to it for streams, events andHIP.device_synchronize().Until now, nonblocking stream synchronization enqueued a host function with
hipLaunchHostFuncthat woke the waiting task through libuv. Because that host function might never run if something went wrong, every synchronization also spawned a second task that polled the stream on a 1 s timer.cooperative_waitpolls the stream first, exactly like before, and then hands a plainhipStreamSynchronizeto a small pool of worker threads while the calling task waits on an event. Errors come back fromhipStreamSynchronizeitself, so the watchdog goes away. This also fixes a few other things:Host functions are slow and get in the way of the stream. CLR puts a blocking marker behind the callback, so later work on the stream waits for the host function too, and the wakeup then goes through the libuv event loop. On CUDA, the equivalent
cuLaunchHostFuncpath added 20–75 µs over a blocking wait, against 2–4 µs for handing the wait to a worker (measured with CUDA.jl on an RTX 5080).Event synchronization blocked the thread.
synchronize(::HIPEvent)polled for a while and then calledhipEventSynchronizeon the calling thread. It now waits on a worker too.HIP.device_synchronize()blocked the thread, and could deadlock with hostcalls. The host side of a hostcall is a task. If a kernel waits for a hostcall whose task can only run on the thread blocked inhipDeviceSynchronize, nothing makes progress. For example, withjulia --threads=1:Every synchronization called
device!from the watchdog task, which also sets the global default device.The workers call HIP directly rather than through the checked wrappers, because those switch to the context of the task making the call. They select the caller's device with
hipSetDevicefirst, since the null stream andhipDeviceSynchronizedepend on it.blocking=trueand thenonblocking_synchronizationpreference work as before. Finalizers, which can't switch tasks, keep the context they selected and wait on the calling thread.This requires GPUToolbox 3.3.1, which fixes two ways the workers could deadlock (JuliaGPU/GPUToolbox.jl#26). Both matter for the hostcall case above.
hipDeviceSynchronizecan't be polled, and before 3.3.1 such a wait had to wait for a free worker when all of them were busy, which never happens if the busy workers are waiting for kernels that need the hostcall to be served. And a worker could run a finalizer that blocks (e.g. freeing memory, which can wait for the device to go idle) before waking up the task it waited for. Now such waits get an extra worker, and workers never run finalizers.This hasn't been run on a discrete AMD GPU yet, so CI is the real validation. I did compare both implementations on the integrated GPU of a Ryzen 9 9950X (gfx1036, run as gfx1030 with ROCm 7.2.4), in one process, alternating between them in blocks. The tables show the median (p90) time for a launch plus
synchronize()of a busy-waiting kernel:--threads=1:--threads=4:Short kernels finish while polling, so they aren't affected. For longer ones, the new version was never slower than the host-function path in any run, and usually had a much tighter tail. The machine was heavily loaded, so the 5 ms numbers varied between runs, but the ordering didn't change.
The trade-off is CPU usage while waiting. HIP busy-waits in
hipStreamSynchronize, so the worker keeps a core busy, just likeblocking=truedoes on the calling thread. On this machine, polling a stream withhipStreamQueryalso kept one of HSA's runtime threads busy, both before and after this PR. As a result, a long wait used about 2 cores instead of about 1.HIP.synchronize(stream; spin=false)skips the polling and brings that back to 1 core, at the cost of a few µs on short kernels.The new tests check that
synchronize(stream)(including the null stream andspin=false),synchronize(event)andHIP.device_synchronize()don't block the thread, and that a hostcall gets served whiledevice_synchronize()waits in a single-threaded process. They don't measure anything. A kernel spins until a flag in host memory is set, and another task on the same thread sets it after yielding many more times than the polling before a wait does, so the wait can only return if that task got to run in the meantime. If it never does, the kernel gives up after 20 s or so and records that it timed out, so a regression fails instead of hanging. That also makes the tests indifferent to the MI250 runner stalling test processes (individual tests there regularly take 5–30× longer than usual, also on main). On main, the event and device tests fail, and the hostcall test times out. With GPUToolbox 3.3.2, they pass on the integrated GPU mentioned above, also with--threads=1and while the process is repeatedly stopped and resumed, and withnonblocking_synchronization=falsethey fail as expected.