Skip to content

POCL: wait for kernels cooperatively - #813

Merged
maleadt merged 4 commits into
mainfrom
tb/pocl-cooperative-wait
Oct 2, 2026
Merged

maleadt merged 4 commits into
mainfrom
tb/pocl-cooperative-wait

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

KernelAbstractions' docs require synchronize to be cooperative: waiting for the device should yield to other tasks instead of blocking the thread inside a C library. The in-tree POCL back-end didn't do this. Because its arrays are plain Arrays that host code can touch at any time, every launch waits for the kernel to finish before returning, and that wait was a blocking clWaitForEvents. So while a POCL kernel ran, nothing else could run on the launching thread: other KA.@spawn tasks, @async I/O, timers.

With this PR, launches still return only once the kernel has completed, but the task yields while waiting. For example, with a kernel that runs for about 400 ms and an @async task on the same thread that ticks every 10 ms:

ticker = @async while !stop[]
    ticks[] += 1
    sleep(0.01)
end
kernel(A, n; ndrange=length(A))

the ticker used to advance 0 times during the launch, and now does 37–38 times.

The waiting itself is done by GPUToolbox.cooperative_wait, the same helper CUDA.jl and OpenCL.jl use. For PoCL, it first polls the kernel's status for at most 10 µs, which catches short kernels. Otherwise PoCL calls us back when the kernel completes (clSetEventCallback), and until then the launching task sleeps while other tasks run.

The approach matters here because PoCL executes kernels on the host's own cores. My first version handed the wait to a worker thread, as CUDA.jl does. That slowed kernels of 0.1–1 ms down by 25–35%, and the benchmark bot showed it: waking the worker just as the kernel starts, and polling for long, both take cores away from the kernel. Callbacks don't need a worker thread of our own (PoCL delivers them from its own callback thread, which costs one extra thread wake-up compared to a blocking wait). GPUToolbox gained support for them in JuliaGPU/GPUToolbox.jl#25.

The wait can't be interrupted: the kernel may be writing to arrays that the caller would otherwise be free to reuse as soon as the launch returns. An interrupt is therefore only rethrown after the kernel has completed.

This adds GPUToolbox as a dependency of KernelAbstractions. It is small (it only depends on LLVM, which KernelAbstractions already uses). Since we depend on it anyway, all calls into PoCL now go through its @gcsafe_ccall, so that blocking in PoCL (compiling programs, waiting for kernels) no longer keeps other threads from collecting garbage. This also matters for the callbacks, which enter Julia from PoCL's own thread.

Performance

Launches of small kernels get faster, since polling notices completion sooner than PoCL wakes up a thread blocked in clWaitForEvents. On the benchmark bot's 4-vCPU runner, saxpy launches of 64–4096 elements went from ~40 µs to ~15 µs.

Longer kernels got 15–30% slower on the bot, though, which was more than the extra wake-up explains. The cause was in GPUToolbox: PoCL's callback thread woke the launching task while holding a spin lock that the task needs right away. With PoCL using every core, the woken thread would often preempt the callback thread on its own CPU and spin until the OS switched back. On 4 cores with 4 PoCL workers, the launching thread spent 195–245 µs of CPU time per ~1 ms launch spinning there, against 20–40 µs with a blocking wait, and that took a core away from the next kernel. JuliaGPU/GPUToolbox.jl#27 fixed this in 3.3.2, bringing it down to 30–65 µs. In an A/B run on 4 cores before that fix, a ~1 ms compute-bound kernel took 1254 µs with callbacks against 1040 µs blocking; with the fix, 1026 µs against 1018 µs. The test machine was shared and heavily loaded, so these are rough numbers; the benchmark bot's next run is the better measurement.

What remains is the extra wake-up for kernels that outlast the poll, a few µs on an idle machine. Removing it would mean having PoCL invoke completion callbacks on the thread that finishes the command, instead of handing them to its callback thread.

Launches that the poll catches allocate as much as before. Ones that wait for the callback allocate 112 bytes more. Which one a launch takes depends on how long the kernel runs, and that made #811's test, which checks that a launch with many arguments allocates no more than one with few, fail randomly on CI. The test now waits by blocking, using an internal switch (POCL.cl.blocking_waits), so it only compares how the arguments are passed.

Testing

A new test checks that a task scheduled before a long POCL launch gets to run during it. It fails with the old blocking wait, and the full test suite passes.

This needs GPUToolbox 3.3.2: 3.3 added the completion callbacks (JuliaGPU/GPUToolbox.jl#25), and 3.3.2 has the fix above (JuliaGPU/GPUToolbox.jl#27). OpenCL.jl gets the same change in JuliaGPU/OpenCL.jl#518.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main c9697db... main / c9697db...
const/@Const/Float32/262144 0.283 ± 0.028 ms 0.297 ± 0.037 ms 0.95 ± 0.15
const/@Const/Float32/65536 0.119 ± 0.025 ms 0.13 ± 0.033 ms 0.92 ± 0.3
const/@Const/Float64/262144 0.425 ± 0.032 ms 0.445 ± 0.04 ms 0.955 ± 0.11
const/@Const/Float64/65536 0.166 ± 0.015 ms 0.187 ± 0.019 ms 0.885 ± 0.12
const/unmarked/Float32/262144 1.13 ± 0.035 ms 1.18 ± 0.036 ms 0.959 ± 0.042
const/unmarked/Float32/65536 0.322 ± 0.034 ms 0.35 ± 0.035 ms 0.92 ± 0.13
const/unmarked/Float64/262144 1.37 ± 0.042 ms 1.41 ± 0.041 ms 0.971 ± 0.041
const/unmarked/Float64/65536 0.387 ± 0.042 ms 0.416 ± 0.037 ms 0.93 ± 0.13
launch/3D static workgroup, dynamic ndrange 0.0758 ± 0.0074 ms 0.0332 ± 0.061 ms 2.28 ± 4.2
launch/3D static workgroup, static ndrange 0.0747 ± 0.02 ms 0.0339 ± 0.061 ms 2.2 ± 4
launch/dynamic workgroup, dynamic ndrange 0.0709 ± 0.033 ms 0.0632 ± 0.061 ms 1.12 ± 1.2
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 0.0697 ± 0.035 ms 0.0606 ± 0.061 ms 1.15 ± 1.3
launch/static workgroup, dynamic ndrange 0.0746 ± 0.0088 ms 0.046 ± 0.061 ms 1.62 ± 2.2
launch/static workgroup, static ndrange 0.0743 ± 0.012 ms 0.0554 ± 0.061 ms 1.34 ± 1.5
partition/dynamic workgroup, dynamic ndrange 0.0597 ± 0.0025 μs 0.0599 ± 0.0017 μs 0.996 ± 0.05
partition/static workgroup, dynamic ndrange 0.0655 ± 0.012 μs 0.068 ± 0.012 μs 0.963 ± 0.25
partition/static workgroup, static ndrange 1.55 ± 0.01 ns 1.55 ± 0.01 ns 1 ± 0.0091
saxpy/default/Float16/1024 0.0745 ± 0.02 ms 0.0559 ± 0.063 ms 1.33 ± 1.5
saxpy/default/Float16/1048576 0.802 ± 0.04 ms 0.833 ± 0.036 ms 0.963 ± 0.063
saxpy/default/Float16/16384 0.0628 ± 0.034 ms 0.0954 ± 0.036 ms 0.659 ± 0.43
saxpy/default/Float16/2048 0.0769 ± 0.013 ms 0.0903 ± 0.044 ms 0.852 ± 0.44
saxpy/default/Float16/256 0.0754 ± 0.0085 ms 0.0547 ± 0.061 ms 1.38 ± 1.5
saxpy/default/Float16/262144 0.247 ± 0.035 ms 0.277 ± 0.036 ms 0.891 ± 0.17
saxpy/default/Float16/32768 0.0724 ± 0.031 ms 0.105 ± 0.037 ms 0.688 ± 0.38
saxpy/default/Float16/4096 0.0758 ± 0.031 ms 0.078 ± 0.03 ms 0.972 ± 0.55
saxpy/default/Float16/512 0.0755 ± 0.012 ms 0.0535 ± 0.061 ms 1.41 ± 1.6
saxpy/default/Float16/64 0.0758 ± 0.0085 ms 0.0513 ± 0.061 ms 1.48 ± 1.8
saxpy/default/Float16/65536 0.0943 ± 0.03 ms 0.128 ± 0.036 ms 0.734 ± 0.31
saxpy/default/Float32/1024 0.0742 ± 0.012 ms 31.4 ± 61 μs 2.36 ± 4.6
saxpy/default/Float32/1048576 0.38 ± 0.033 ms 0.424 ± 0.048 ms 0.897 ± 0.13
saxpy/default/Float32/16384 0.0564 ± 0.032 ms 0.0761 ± 0.037 ms 0.741 ± 0.55
saxpy/default/Float32/2048 0.0743 ± 0.022 ms 0.0627 ± 0.063 ms 1.19 ± 1.2
saxpy/default/Float32/256 0.0755 ± 0.016 ms 0.0426 ± 0.061 ms 1.77 ± 2.5
saxpy/default/Float32/262144 0.144 ± 0.028 ms 0.17 ± 0.034 ms 0.849 ± 0.24
saxpy/default/Float32/32768 0.0613 ± 0.031 ms 0.0907 ± 0.036 ms 0.675 ± 0.43
saxpy/default/Float32/4096 0.0748 ± 0.028 ms 0.0726 ± 0.057 ms 1.03 ± 0.9
saxpy/default/Float32/512 0.0747 ± 0.012 ms 30.9 ± 61 μs 2.42 ± 4.8
saxpy/default/Float32/64 0.0754 ± 0.011 ms 0.0456 ± 0.06 ms 1.65 ± 2.2
saxpy/default/Float32/65536 0.0745 ± 0.03 ms 0.102 ± 0.035 ms 0.729 ± 0.38
saxpy/default/Float64/1024 0.074 ± 0.012 ms 0.0417 ± 0.064 ms 1.77 ± 2.7
saxpy/default/Float64/1048576 0.608 ± 0.074 ms 0.689 ± 0.08 ms 0.882 ± 0.15
saxpy/default/Float64/16384 0.0561 ± 0.029 ms 0.0919 ± 0.037 ms 0.611 ± 0.4
saxpy/default/Float64/2048 0.0709 ± 0.03 ms 0.0673 ± 0.065 ms 1.05 ± 1.1
saxpy/default/Float64/256 0.0751 ± 0.011 ms 30.8 ± 61 μs 2.44 ± 4.8
saxpy/default/Float64/262144 0.198 ± 0.039 ms 0.221 ± 0.04 ms 0.897 ± 0.24
saxpy/default/Float64/32768 0.0687 ± 0.031 ms 0.104 ± 0.037 ms 0.661 ± 0.38
saxpy/default/Float64/4096 0.069 ± 0.028 ms 0.0706 ± 0.06 ms 0.977 ± 0.92
saxpy/default/Float64/512 0.0743 ± 0.011 ms 31 ± 62 μs 2.4 ± 4.8
saxpy/default/Float64/64 0.0759 ± 0.013 ms 0.0476 ± 0.061 ms 1.59 ± 2
saxpy/default/Float64/65536 0.09 ± 0.032 ms 0.121 ± 0.035 ms 0.741 ± 0.34
saxpy/static workgroup=(1024,)/Float16/1024 0.0745 ± 0.011 ms 0.0517 ± 0.062 ms 1.44 ± 1.8
saxpy/static workgroup=(1024,)/Float16/1048576 0.8 ± 0.042 ms 0.83 ± 0.038 ms 0.965 ± 0.067
saxpy/static workgroup=(1024,)/Float16/16384 0.062 ± 0.032 ms 0.0853 ± 0.034 ms 0.726 ± 0.47
saxpy/static workgroup=(1024,)/Float16/2048 0.0768 ± 0.012 ms 0.0786 ± 0.049 ms 0.977 ± 0.63
saxpy/static workgroup=(1024,)/Float16/256 0.0747 ± 0.024 ms 0.0543 ± 0.061 ms 1.37 ± 1.6
saxpy/static workgroup=(1024,)/Float16/262144 0.248 ± 0.035 ms 0.274 ± 0.034 ms 0.906 ± 0.17
saxpy/static workgroup=(1024,)/Float16/32768 0.0727 ± 0.03 ms 0.101 ± 0.035 ms 0.723 ± 0.39
saxpy/static workgroup=(1024,)/Float16/4096 0.0712 ± 0.032 ms 0.0914 ± 0.032 ms 0.779 ± 0.44
saxpy/static workgroup=(1024,)/Float16/512 0.0749 ± 0.017 ms 0.0611 ± 0.061 ms 1.23 ± 1.3
saxpy/static workgroup=(1024,)/Float16/64 0.0749 ± 0.012 ms 0.0575 ± 0.06 ms 1.3 ± 1.4
saxpy/static workgroup=(1024,)/Float16/65536 0.0952 ± 0.029 ms 0.123 ± 0.034 ms 0.772 ± 0.32
saxpy/static workgroup=(1024,)/Float32/1024 0.074 ± 0.011 ms 0.0324 ± 0.062 ms 2.28 ± 4.4
saxpy/static workgroup=(1024,)/Float32/1048576 0.404 ± 0.037 ms 0.436 ± 0.045 ms 0.926 ± 0.13
saxpy/static workgroup=(1024,)/Float32/16384 0.0575 ± 0.029 ms 0.0806 ± 0.039 ms 0.713 ± 0.5
saxpy/static workgroup=(1024,)/Float32/2048 0.0744 ± 0.013 ms 0.0625 ± 0.063 ms 1.19 ± 1.2
saxpy/static workgroup=(1024,)/Float32/256 0.0747 ± 0.014 ms 0.0583 ± 0.061 ms 1.28 ± 1.4
saxpy/static workgroup=(1024,)/Float32/262144 0.152 ± 0.029 ms 0.172 ± 0.03 ms 0.881 ± 0.23
saxpy/static workgroup=(1024,)/Float32/32768 0.0626 ± 0.028 ms 0.089 ± 0.033 ms 0.703 ± 0.41
saxpy/static workgroup=(1024,)/Float32/4096 0.0759 ± 0.027 ms 0.0745 ± 0.066 ms 1.02 ± 0.98
saxpy/static workgroup=(1024,)/Float32/512 0.074 ± 0.025 ms 0.0329 ± 0.061 ms 2.25 ± 4.3
saxpy/static workgroup=(1024,)/Float32/64 0.0747 ± 0.018 ms 0.0647 ± 0.061 ms 1.15 ± 1.1
saxpy/static workgroup=(1024,)/Float32/65536 0.0774 ± 0.03 ms 0.105 ± 0.032 ms 0.739 ± 0.36
saxpy/static workgroup=(1024,)/Float64/1024 0.0742 ± 0.0067 ms 0.0357 ± 0.063 ms 2.08 ± 3.7
saxpy/static workgroup=(1024,)/Float64/1048576 0.575 ± 0.075 ms 0.647 ± 0.092 ms 0.889 ± 0.17
saxpy/static workgroup=(1024,)/Float64/16384 0.0604 ± 0.028 ms 0.0824 ± 0.036 ms 0.733 ± 0.47
saxpy/static workgroup=(1024,)/Float64/2048 0.0667 ± 0.03 ms 0.0602 ± 0.064 ms 1.11 ± 1.3
saxpy/static workgroup=(1024,)/Float64/256 0.0749 ± 0.017 ms 0.0398 ± 0.061 ms 1.88 ± 2.9
saxpy/static workgroup=(1024,)/Float64/262144 0.202 ± 0.036 ms 0.221 ± 0.041 ms 0.911 ± 0.23
saxpy/static workgroup=(1024,)/Float64/32768 0.0737 ± 0.029 ms 0.0991 ± 0.033 ms 0.743 ± 0.38
saxpy/static workgroup=(1024,)/Float64/4096 0.0646 ± 0.027 ms 0.0646 ± 0.068 ms 1 ± 1.1
saxpy/static workgroup=(1024,)/Float64/512 0.074 ± 0.011 ms 0.032 ± 0.062 ms 2.31 ± 4.5
saxpy/static workgroup=(1024,)/Float64/64 0.0751 ± 0.0095 ms 0.0608 ± 0.06 ms 1.24 ± 1.2
saxpy/static workgroup=(1024,)/Float64/65536 0.094 ± 0.032 ms 0.118 ± 0.031 ms 0.798 ± 0.34
time_to_load 0.807 ± 0.0097 s 0.837 ± 0.015 s 0.964 ± 0.021
main c9697db... main / c9697db...
const/@Const/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float64/262144 2 allocs: 32 B 6 allocs: 0.141 kB 0.222
const/unmarked/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
launch/3D static workgroup, dynamic ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/3D static workgroup, static ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/dynamic workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, static ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/default/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/32768 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/default/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/default/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/default/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/262144 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/default/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/static workgroup=(1024,)/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/static workgroup=(1024,)/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/262144 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/static workgroup=(1024,)/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/static workgroup=(1024,)/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/262144 5 allocs: 0.0781 kB 9 allocs: 0.188 kB 0.417
saxpy/static workgroup=(1024,)/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

Calls into PoCL can block for a long time: building programs, and waiting for
kernels, which the POCL back-end does on every launch. As plain ccalls, they
prevent every other thread from collecting garbage in the meantime.

Use GPUToolbox's `@gcsafe_ccall` for all of them, as CUDA.jl and OpenCL.jl do.
KernelAbstractions requires `synchronize` to be cooperative: waiting for the
device should yield to other tasks rather than block the thread. The POCL
back-end launches synchronously, because its arrays are plain `Array`s, but
every launch waited for its kernel with a blocking `clWaitForEvents`, so no
other task could run on the launching thread in the meantime.

Wait using GPUToolbox's `cooperative_wait` instead: poll the kernel's status for
at most 10 µs, and then have PoCL notify us of its completion through
`clSetEventCallback`. Waiting on a worker thread, or polling for longer, would
compete with the kernel for the CPU cores it executes on, making it
considerably slower. The wait cannot be interrupted, as the kernel may be using
memory the caller would release: interrupts are only thrown once the kernel
completed.

The first launch after new code has been defined allocates to look up the
completion callback again, so measure allocations from a warmed-up function.
Whether a launch allocates now depends on how it waits: one that polling finds
completed doesn't, one that waits for the completion callback does. Which one a
launch takes depends on how long the kernel runs, so the test comparing a launch
with many arguments to one with few failed randomly on CI (e.g. 144 > 32 bytes).

Add an internal switch to wait for kernels by blocking instead, and use it in
that test, so that it only compares how the arguments are passed.
3.3.2 keeps the task woken by a completion callback from spinning on a lock
held by PoCL's callback thread after preempting it. With PoCL using every core,
that took a core away from the next kernel, making 50-400 µs kernels 15-30%
slower on the benchmark bot (JuliaGPU/GPUToolbox.jl#27).
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 82.35294% with 9 lines in your changes missing coverage. Please review.
✅ Project coverage is 69.82%. Comparing base (225a67a) to head (c9697db).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/pocl/nanoOpenCL.jl 80.85% 9 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #813      +/-   ##
==========================================
+ Coverage   69.54%   69.82%   +0.27%     
==========================================
  Files          26       26              
  Lines        2157     2177      +20     
==========================================
+ Hits         1500     1520      +20     
  Misses        657      657              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt

maleadt commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

I reran all 83 benchmarks against the merge base and PR head (Julia 1.13, PoCL 7.2.0+21, GPUToolbox 3.3.2; one Julia thread and four PoCL workers pinned to four CPUs). All 54 saxpy cases through 65k elements improved (~19–22% on average), but all six 262k cases regressed 10–32%; 1M cases were mixed. Launch improved ~19%; const and partition were essentially unchanged. For example, Float16 saxpy at 262k went from 36.55 to 44.68 µs, and Float32 from 56.29 to 69.02 µs.

A minimal 262k Float16 reproducer, alternating waits in the same PR process, measured 36.28–38.17 µs blocking versus 47.65–47.99 µs cooperative. Raw PoCL events reproduce the gap, including with spin=false: the cooperative wait used ~37–42 µs of Julia-thread CPU during a ~37–42 µs wait, versus ~0.9 µs of CPU during a ~29–33 µs blocking wait. Giving Julia a fifth CPU substantially reduced the gap. This points to the Julia thread remaining active during the task-level wait and contending with PoCL's CPU workers, not the explicit 10 µs spin or callback latency. Host load makes exact ratios noisy, but the medium-kernel regression is reproducible.

TL;DR: this improves latency, but because it allows the Julia scheduler to be active during kernel execution (even if only running the idle loop), it slightly lowers performance on CPU-constrained workers like on CI. This is less visible on actual GPU back-ends because then we aren't contending with kernels executing on CPUs there. I don't think there's much we can do about this, so I'll go ahead and merge.

@maleadt
maleadt merged commit 9bc9560 into main Oct 2, 2026
78 of 79 checks passed
@maleadt
maleadt deleted the tb/pocl-cooperative-wait branch October 2, 2026 08:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants