Don't let tasks woken by completion callbacks spin on a preempted lock holder - #27
Merged
Merged
Conversation
A driver thread signals a completion by taking its spin lock and waking the waiting task, which then takes the lock again. If the OS runs that task on the driver thread's CPU, it preempts the driver thread while that still holds the lock, and spins until the OS switches back. When all cores are busy, that can take hundreds of µs: with KernelAbstractions' PoCL back-end on 4 cores, the launching thread spent 200-250 µs of CPU time per launch of a ~1 ms kernel spinning there (instead of 35-40 µs with a blocking wait), taking a core away from the next kernel. Use a spin lock whose waiters give up their CPU (`sched_yield`) after spinning briefly, so that the preempted driver thread gets to release the lock.
maleadt
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 1, 2026
3.3.2 keeps the task woken by a completion callback from spinning on a lock held by PoCL's callback thread after preempting it. With PoCL using every core, that took a core away from the next kernel, making 50-400 µs kernels 15-30% slower on the benchmark bot (JuliaGPU/GPUToolbox.jl#27).
maleadt
added a commit
to JuliaGPU/OpenCL.jl
that referenced
this pull request
Oct 1, 2026
On CPU devices, which wait for completion callbacks, 3.3.2 keeps the woken task from spinning on a lock held by the driver's callback thread after preempting it, which made longer PoCL kernels slower when all cores are busy (JuliaGPU/GPUToolbox.jl#27).
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
KernelAbstractions' benchmark bot still showed PoCL launches of 50–400 µs kernels getting 15–30% slower with cooperative waiting (JuliaGPU/KernelAbstractions.jl#813), much more than the one extra thread wake-up that completion callbacks cost. The cause turned out to be in
signal_completion, and it's a classic lock-holder preemption problem.signal_completionruns on the driver's thread (for PoCL, its dedicated callback thread). It takes the completion's spin lock, wakes the waiting task withnotify, and then releases the lock. The woken task immediately takes the lock again (waiton a condition reacquires it before returning). If the OS runs the woken Julia thread on the same CPU as the driver's thread, the Julia thread preempts it while it still holds the lock. Julia'sSpinLocknever gives up the CPU, so the task spins until the OS happens to switch back. That's rare on an idle machine with spare cores. But PoCL runs kernels on the host's cores, and with as many PoCL workers as cores, a GitHub runner's 4 vCPUs for example, there's never a spare one.perfon the launching thread, with KernelAbstractions launching ~1 ms kernels in a loop on 4 cores (POCL_MAX_PTHREAD_COUNT=4, pinned withtaskset), shows where its time goes:Measured as the launching thread's CPU time per launch, which a loaded machine doesn't distort much:
clWaitForEventscooperative_waitwith callbacks, GPUToolbox 3.3.1That spinning takes a core away from the next kernel, which is where the slowdown came from. In one A/B run on 4 cores, a 1M-element saxpy took 108 µs with blocking waits and 194 µs with callbacks, and a ~1 ms compute-bound kernel 1040 µs versus 1254 µs. With this PR, the compute-bound kernel took 1026 µs against 1018 µs blocking. (The machine was shared and heavily loaded, so only compare numbers within one run.)
The fix is a spin lock whose waiters give up their CPU (
sched_yield, orSwitchToThreadon Windows) after spinning briefly, so the preempted driver thread gets to run and release the lock. It wrapsThreads.SpinLock, and the completion uses it through a plainGenericCondition, so the wake-up protocol, including the handling of interrupts, is the same as before.I first tried keeping the spin lock and waking the task only after releasing it. That fixes the problem too, but needs
GenericConditioninternals: its wait queue holds tasks up to 1.13, andBase.WaitEntry1objects on nightly. It also opens a small window in which a waiter that is interrupted gets scheduled again later, spuriously.The new tests let a foreign thread hold the lock (sleeping) after notifying the waiter, and contend for it while a Julia thread holds it and runs the GC. They check that both sides finish, rather than timing anything. The test suite passes on Julia 1.10, 1.12, 1.13 and nightly. sol reviewed the lock implementation (
AbstractLockfallbacks forGenericConditionfrom 1.10 to nightly, finalizer inhibition and GC safepoints on a foreign thread, the Windowsccall).This bumps the version to 3.3.2.