POCL: wait for kernels cooperatively - #813
Conversation
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Calls into PoCL can block for a long time: building programs, and waiting for kernels, which the POCL back-end does on every launch. As plain ccalls, they prevent every other thread from collecting garbage in the meantime. Use GPUToolbox's `@gcsafe_ccall` for all of them, as CUDA.jl and OpenCL.jl do.
43649ac to
b1a6b02
Compare
KernelAbstractions requires `synchronize` to be cooperative: waiting for the device should yield to other tasks rather than block the thread. The POCL back-end launches synchronously, because its arrays are plain `Array`s, but every launch waited for its kernel with a blocking `clWaitForEvents`, so no other task could run on the launching thread in the meantime. Wait using GPUToolbox's `cooperative_wait` instead: poll the kernel's status for at most 10 µs, and then have PoCL notify us of its completion through `clSetEventCallback`. Waiting on a worker thread, or polling for longer, would compete with the kernel for the CPU cores it executes on, making it considerably slower. The wait cannot be interrupted, as the kernel may be using memory the caller would release: interrupts are only thrown once the kernel completed. The first launch after new code has been defined allocates to look up the completion callback again, so measure allocations from a warmed-up function.
b1a6b02 to
6ad5ab3
Compare
Whether a launch allocates now depends on how it waits: one that polling finds completed doesn't, one that waits for the completion callback does. Which one a launch takes depends on how long the kernel runs, so the test comparing a launch with many arguments to one with few failed randomly on CI (e.g. 144 > 32 bytes). Add an internal switch to wait for kernels by blocking instead, and use it in that test, so that it only compares how the arguments are passed.
3.3.2 keeps the task woken by a completion callback from spinning on a lock held by PoCL's callback thread after preempting it. With PoCL using every core, that took a core away from the next kernel, making 50-400 µs kernels 15-30% slower on the benchmark bot (JuliaGPU/GPUToolbox.jl#27).
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #813 +/- ##
==========================================
+ Coverage 69.54% 69.82% +0.27%
==========================================
Files 26 26
Lines 2157 2177 +20
==========================================
+ Hits 1500 1520 +20
Misses 657 657 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
TL;DR: this improves latency, but because it allows the Julia scheduler to be active during kernel execution (even if only running the idle loop), it slightly lowers performance on CPU-constrained workers like on CI. This is less visible on actual GPU back-ends because then we aren't contending with kernels executing on CPUs there. I don't think there's much we can do about this, so I'll go ahead and merge. |
KernelAbstractions' docs require
synchronizeto be cooperative: waiting for the device should yield to other tasks instead of blocking the thread inside a C library. The in-tree POCL back-end didn't do this. Because its arrays are plainArrays that host code can touch at any time, every launch waits for the kernel to finish before returning, and that wait was a blockingclWaitForEvents. So while a POCL kernel ran, nothing else could run on the launching thread: otherKA.@spawntasks,@asyncI/O, timers.With this PR, launches still return only once the kernel has completed, but the task yields while waiting. For example, with a kernel that runs for about 400 ms and an
@asynctask on the same thread that ticks every 10 ms:the ticker used to advance 0 times during the launch, and now does 37–38 times.
The waiting itself is done by
GPUToolbox.cooperative_wait, the same helper CUDA.jl and OpenCL.jl use. For PoCL, it first polls the kernel's status for at most 10 µs, which catches short kernels. Otherwise PoCL calls us back when the kernel completes (clSetEventCallback), and until then the launching task sleeps while other tasks run.The approach matters here because PoCL executes kernels on the host's own cores. My first version handed the wait to a worker thread, as CUDA.jl does. That slowed kernels of 0.1–1 ms down by 25–35%, and the benchmark bot showed it: waking the worker just as the kernel starts, and polling for long, both take cores away from the kernel. Callbacks don't need a worker thread of our own (PoCL delivers them from its own callback thread, which costs one extra thread wake-up compared to a blocking wait). GPUToolbox gained support for them in JuliaGPU/GPUToolbox.jl#25.
The wait can't be interrupted: the kernel may be writing to arrays that the caller would otherwise be free to reuse as soon as the launch returns. An interrupt is therefore only rethrown after the kernel has completed.
This adds GPUToolbox as a dependency of KernelAbstractions. It is small (it only depends on LLVM, which KernelAbstractions already uses). Since we depend on it anyway, all calls into PoCL now go through its
@gcsafe_ccall, so that blocking in PoCL (compiling programs, waiting for kernels) no longer keeps other threads from collecting garbage. This also matters for the callbacks, which enter Julia from PoCL's own thread.Performance
Launches of small kernels get faster, since polling notices completion sooner than PoCL wakes up a thread blocked in
clWaitForEvents. On the benchmark bot's 4-vCPU runner, saxpy launches of 64–4096 elements went from ~40 µs to ~15 µs.Longer kernels got 15–30% slower on the bot, though, which was more than the extra wake-up explains. The cause was in GPUToolbox: PoCL's callback thread woke the launching task while holding a spin lock that the task needs right away. With PoCL using every core, the woken thread would often preempt the callback thread on its own CPU and spin until the OS switched back. On 4 cores with 4 PoCL workers, the launching thread spent 195–245 µs of CPU time per ~1 ms launch spinning there, against 20–40 µs with a blocking wait, and that took a core away from the next kernel. JuliaGPU/GPUToolbox.jl#27 fixed this in 3.3.2, bringing it down to 30–65 µs. In an A/B run on 4 cores before that fix, a ~1 ms compute-bound kernel took 1254 µs with callbacks against 1040 µs blocking; with the fix, 1026 µs against 1018 µs. The test machine was shared and heavily loaded, so these are rough numbers; the benchmark bot's next run is the better measurement.
What remains is the extra wake-up for kernels that outlast the poll, a few µs on an idle machine. Removing it would mean having PoCL invoke completion callbacks on the thread that finishes the command, instead of handing them to its callback thread.
Launches that the poll catches allocate as much as before. Ones that wait for the callback allocate 112 bytes more. Which one a launch takes depends on how long the kernel runs, and that made #811's test, which checks that a launch with many arguments allocates no more than one with few, fail randomly on CI. The test now waits by blocking, using an internal switch (
POCL.cl.blocking_waits), so it only compares how the arguments are passed.Testing
A new test checks that a task scheduled before a long POCL launch gets to run during it. It fails with the old blocking wait, and the full test suite passes.
This needs GPUToolbox 3.3.2: 3.3 added the completion callbacks (JuliaGPU/GPUToolbox.jl#25), and 3.3.2 has the fix above (JuliaGPU/GPUToolbox.jl#27). OpenCL.jl gets the same change in JuliaGPU/OpenCL.jl#518.