KernelInterface 0.4: launch kernels with many arguments without allocating - #811
Merged
Merged
Conversation
Each layer of PoCL's launch path took the kernel arguments as varargs and splatted them into the next: the kernel object's call, `clcall`, `call`, and `set_args!`, which recursed with one splat per argument. Julia doesn't turn a splat of more than 32 elements into a direct call, and a method with both varargs and keyword arguments splats them into its body, so launching a kernel with many arguments went through `Core._apply_iterate` several times. Pass the arguments along as one tuple instead, and generate the per-argument code. The kernel object keeps its call syntax, with its keyword method defined explicitly, and `launch_tuple` takes the tuple directly. `@opencl` builds the argument types without `map`, which isn't type stable for 32 or more elements.
The back-end hook was `launch(kernel, groups, items, args...; kwargs...)`.
A method with both varargs and keyword arguments splats the arguments into its
body, and Julia doesn't turn a splat of more than 32 elements into a direct
call, so a launch with many arguments allocated in every back end, whatever the
caller did. Avoiding that meant writing a `Core.kwcall` method by hand in each
back end.
Make the hook take the arguments as one tuple instead:
launch(kernel::Kernel{<:NewBackend}, groups::Dims{3}, items::Dims{3}, args::Tuple; kwargs...)
A back end can then pass the tuple on to its own launcher. This is breaking, so
this is KernelInterface 0.4. No back end depends on 0.3 yet.
Calling a `Kernel` keeps its syntax. Its keyword method is defined explicitly,
so that it can pass the arguments on as a tuple, and `@launch` builds the
argument types without `map`, which isn't type stable for 32 or more elements.
Member
Author
|
Could make this 0.3.1 since no back-end has been adapted yet, but it doesn't hurt being semver-strict. |
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
The generic launch splatted the kernel arguments from the `Kernel` call down to the `KI.Kernel` call, and mapped over them to compute the argument types. With more than 32 arguments neither is a direct call, so on PoCL a kernel with 40 arguments allocated 25.7 KB per launch, against 32 bytes with 4 arguments. Pass them along as a tuple, as `KI.Kernel` now does, and generate the code that computes the argument types and calls the `KI.Kernel`. A 40-argument launch now allocates as much as a 4-argument one. CUDA.jl#3309 did the same for CUDA's own launch of KA kernels, which the generic launch replaces.
Converting an `Array` argument for PoCL yields a device array that only holds a pointer, and nothing kept the original arguments alive after they were converted. The garbage collector could therefore free an array while its remaining arguments were converted, or while the kernel ran. A launch through KernelAbstractions happens to keep them rooted, but calling a PoCL kernel directly doesn't: with a forced collection during argument conversion, the array passed to the kernel was finalized before the launch was queued. Preserve the arguments while converting them and queuing the launch, and in `KI.launch`, which waits for the kernel, until it completes. The latter matters once waiting no longer blocks the garbage collector.
maleadt
force-pushed
the
tb/many-args
branch
from
September 30, 2026 16:48
e8376e8 to
c26da5a
Compare
This was referenced Sep 30, 2026
`KI.@launch` and KernelAbstractions' generic launch converted the callable with `argconvert` before passing it to `kernel_function`. For a closure, the converted callable only holds pointers to the arrays it captures, so nothing kept those arrays alive: a kernel compiled with `launch=false` could read freed memory once the garbage collector had run. A back end also couldn't see the captured arrays at launch, which Metal needs to declare them to the command encoder, as `@metal` has done since Metal.jl#992. `kernel_function` now receives the original callable and converts it itself, and the kernel keeps the original alive. PoCL's kernels keep it next to the compiled kernel, and root it while launching. KernelInterface's testsuite checks that a kernel whose closure captures an array still works after a collection.
This was referenced Sep 30, 2026
Intel's CPU runtime skips the partial last work-group along x in N-d launches
JuliaGPU/OpenCL.jl#521
Open
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #811 +/- ##
==========================================
+ Coverage 69.46% 69.54% +0.07%
==========================================
Files 26 26
Lines 2142 2157 +15
==========================================
+ Hits 1488 1500 +12
- Misses 654 657 +3 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A
@kernelwith many arguments is expensive to launch. With 40 arguments, a launch on the CPU back-end allocates 25.7 KB, against 32 bytes with 4:The cause is three limits in Julia:
mapover 32 or more elements isn't type stable;The launch path hits all three: in KA's generic launch, in KernelInterface's
Kernelcall and itsKI.launchhook, and in several layers of the PoCL back-end.CUDA.jl fixed this for its own launch path in JuliaGPU/CUDA.jl#3309, which includes launching KA kernels. The generic launch from #801 will replace that code in CUDA.jl, so porting CUDA to it would bring the problem back without this PR.
This PR passes the arguments along as one tuple, from the call where they enter to the back end's own launcher. Where code has to handle each argument (computing the argument types, setting the kernel arguments), it is generated instead of mapping or splatting. The public calls keep their syntax. Their keyword methods (
Core.kwcall) are defined explicitly, as Base does forinvokelatest, so that they can forward the tuple. As in CUDA.jl, the downside is thathasmethodand REPL completion no longer see the keyword names of these calls.The first breaking change is KernelInterface's back-end hook, which now receives the tuple:
With the varargs signature, every back end would have had to write its own
Core.kwcallmethod to avoid the splat. No back end depends on KernelInterface 0.3 yet: 0.3.0 was registered today, and the back-end ports aren't merged. So KernelInterface on main becomes 0.4.0-dev, KA 0.10 will require it, and the ports implement the new hook directly.The PR also fixes an older bug in the PoCL back-end. It converts
Arrayarguments to device arrays that hold only a pointer, and nothing kept the original arrays alive. Launching through KA happened to keep them rooted, but calling a PoCL kernel directly didn't, and with a garbage collection during argument conversion, the array was freed before the launch was queued. The arguments are now kept alive until the kernel has completed. That will matter once waiting for a PoCL kernel no longer blocks the garbage collector.A second change to the back-end contract fixes a bug in how KernelInterface compiles closures.
KI.@launch, and KA's generic launch, converted the callable withargconvertbefore passing it tokernel_function. The converted form of a closure only holds pointers to the arrays it captures, so nothing kept those arrays alive, and a kernel compiled withlaunch=falsecould read freed memory once the garbage collector had run:A back end also couldn't see the captured arrays at launch, which Metal needs, to declare them to the command encoder (Metal.jl#992 fixed the same for
@metal).kernel_functionnow receives the original callable and converts it itself, and the kernel it returns keeps the original alive. A back end can convert it again at every launch if it needs to. KernelInterface's testsuite checks this with a kernel whose closure captures an array.Per launch on PoCL, with Julia 1.13 and
POCL_MAX_PTHREAD_COUNT=1:The launch testsuite that back ends run now includes a kernel with 40 arguments, and the PoCL tests check that launching it allocates no more than launching one with 4. The KernelInterface tests check that the arguments reach
KI.launchas one tuple, including a single tuple-valued argument. Tested on Julia 1.10 and 1.13.