Skip to content

Don't let tasks woken by completion callbacks spin on a preempted lock holder - #27

Merged
maleadt merged 2 commits into
mainfrom
tb/signal-unlocked
Oct 1, 2026
Merged

maleadt merged 2 commits into
mainfrom
tb/signal-unlocked

Conversation

@maleadt

@maleadt maleadt commented Oct 1, 2026

Copy link
Copy Markdown
Member

KernelAbstractions' benchmark bot still showed PoCL launches of 50–400 µs kernels getting 15–30% slower with cooperative waiting (JuliaGPU/KernelAbstractions.jl#813), much more than the one extra thread wake-up that completion callbacks cost. The cause turned out to be in signal_completion, and it's a classic lock-holder preemption problem.

signal_completion runs on the driver's thread (for PoCL, its dedicated callback thread). It takes the completion's spin lock, wakes the waiting task with notify, and then releases the lock. The woken task immediately takes the lock again (wait on a condition reacquires it before returning). If the OS runs the woken Julia thread on the same CPU as the driver's thread, the Julia thread preempts it while it still holds the lock. Julia's SpinLock never gives up the CPU, so the task spins until the OS happens to switch back. That's rare on an idle machine with spare cores. But PoCL runs kernels on the host's cores, and with as many PoCL workers as cores, a GitHub runner's 4 vCPUs for example, there's never a spare one.

perf on the launching thread, with KernelAbstractions launching ~1 ms kernels in a loop on 4 cores (POCL_MAX_PTHREAD_COUNT=4, pinned with taskset), shows where its time goes:

82.03%  julia  sys.so  [.] julia_lock_1409.2
        |--75.65%--lock (inlined)
        |          relockall; (inlined)
        |          #wait#392 (inlined)

Measured as the launching thread's CPU time per launch, which a loaded machine doesn't distort much:

~1 ms kernel, 4 PoCL workers on 4 cores launching thread's CPU per launch
blocking clWaitForEvents 20–40 µs
cooperative_wait with callbacks, GPUToolbox 3.3.1 195–245 µs
same, with this PR 30–65 µs

That spinning takes a core away from the next kernel, which is where the slowdown came from. In one A/B run on 4 cores, a 1M-element saxpy took 108 µs with blocking waits and 194 µs with callbacks, and a ~1 ms compute-bound kernel 1040 µs versus 1254 µs. With this PR, the compute-bound kernel took 1026 µs against 1018 µs blocking. (The machine was shared and heavily loaded, so only compare numbers within one run.)

The fix is a spin lock whose waiters give up their CPU (sched_yield, or SwitchToThread on Windows) after spinning briefly, so the preempted driver thread gets to run and release the lock. It wraps Threads.SpinLock, and the completion uses it through a plain GenericCondition, so the wake-up protocol, including the handling of interrupts, is the same as before.

I first tried keeping the spin lock and waking the task only after releasing it. That fixes the problem too, but needs GenericCondition internals: its wait queue holds tasks up to 1.13, and Base.WaitEntry1 objects on nightly. It also opens a small window in which a waiter that is interrupted gets scheduled again later, spuriously.

The new tests let a foreign thread hold the lock (sleeping) after notifying the waiter, and contend for it while a Julia thread holds it and runs the GC. They check that both sides finish, rather than timing anything. The test suite passes on Julia 1.10, 1.12, 1.13 and nightly. sol reviewed the lock implementation (AbstractLock fallbacks for GenericCondition from 1.10 to nightly, finalizer inhibition and GC safepoints on a foreign thread, the Windows ccall).

This bumps the version to 3.3.2.

A driver thread signals a completion by taking its spin lock and waking the
waiting task, which then takes the lock again. If the OS runs that task on the
driver thread's CPU, it preempts the driver thread while that still holds the
lock, and spins until the OS switches back. When all cores are busy, that can
take hundreds of µs: with KernelAbstractions' PoCL back-end on 4 cores, the
launching thread spent 200-250 µs of CPU time per launch of a ~1 ms kernel
spinning there (instead of 35-40 µs with a blocking wait), taking a core away
from the next kernel.

Use a spin lock whose waiters give up their CPU (`sched_yield`) after spinning
briefly, so that the preempted driver thread gets to release the lock.
@maleadt
maleadt merged commit 46c9375 into main Oct 1, 2026
9 of 10 checks passed
@maleadt
maleadt deleted the tb/signal-unlocked branch October 1, 2026 12:45
maleadt added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 1, 2026
3.3.2 keeps the task woken by a completion callback from spinning on a lock
held by PoCL's callback thread after preempting it. With PoCL using every core,
that took a core away from the next kernel, making 50-400 µs kernels 15-30%
slower on the benchmark bot (JuliaGPU/GPUToolbox.jl#27).
maleadt added a commit to JuliaGPU/OpenCL.jl that referenced this pull request Oct 1, 2026
On CPU devices, which wait for completion callbacks, 3.3.2 keeps the woken
task from spinning on a lock held by the driver's callback thread after
preempting it, which made longer PoCL kernels slower when all cores are busy
(JuliaGPU/GPUToolbox.jl#27).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant