Skip to content

KernelInterface 0.4: launch kernels with many arguments without allocating - #811

Merged
maleadt merged 5 commits into
mainfrom
tb/many-args
Sep 30, 2026
Merged

maleadt merged 5 commits into
mainfrom
tb/many-args

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

A @kernel with many arguments is expensive to launch. With 40 arguments, a launch on the CPU back-end allocates 25.7 KB, against 32 bytes with 4:

xs = [Symbol(:x, i) for i in 1:40]
@eval @kernel function many!(A, $(xs...))
    I = @index(Global, Linear)
    @inbounds A[I] = $(foldl((a, b) -> :($a + $b), xs))
end

A = zeros(Int, 64)
k = many!(CPU(), 64)
@eval launch() = k(A, $((1:40)...); ndrange = 64)
launch()
@allocated launch()   # 25680 on main, 32 with this PR

The cause is three limits in Julia:

  • splatting more than 32 elements never becomes a direct call;
  • map over 32 or more elements isn't type stable;
  • a method that takes both varargs and keyword arguments splats its arguments to call its body.

The launch path hits all three: in KA's generic launch, in KernelInterface's Kernel call and its KI.launch hook, and in several layers of the PoCL back-end.

CUDA.jl fixed this for its own launch path in JuliaGPU/CUDA.jl#3309, which includes launching KA kernels. The generic launch from #801 will replace that code in CUDA.jl, so porting CUDA to it would bring the problem back without this PR.

This PR passes the arguments along as one tuple, from the call where they enter to the back end's own launcher. Where code has to handle each argument (computing the argument types, setting the kernel arguments), it is generated instead of mapping or splatting. The public calls keep their syntax. Their keyword methods (Core.kwcall) are defined explicitly, as Base does for invokelatest, so that they can forward the tuple. As in CUDA.jl, the downside is that hasmethod and REPL completion no longer see the keyword names of these calls.

The first breaking change is KernelInterface's back-end hook, which now receives the tuple:

# KernelInterface 0.3
KI.launch(kernel::KI.Kernel{B}, groups::Dims{3}, items::Dims{3}, args::Vararg{Any, N}; kwargs...) where {N}
# KernelInterface 0.4
KI.launch(kernel::KI.Kernel{B}, groups::Dims{3}, items::Dims{3}, args::Tuple; kwargs...)

With the varargs signature, every back end would have had to write its own Core.kwcall method to avoid the splat. No back end depends on KernelInterface 0.3 yet: 0.3.0 was registered today, and the back-end ports aren't merged. So KernelInterface on main becomes 0.4.0-dev, KA 0.10 will require it, and the ports implement the new hook directly.

The PR also fixes an older bug in the PoCL back-end. It converts Array arguments to device arrays that hold only a pointer, and nothing kept the original arrays alive. Launching through KA happened to keep them rooted, but calling a PoCL kernel directly didn't, and with a garbage collection during argument conversion, the array was freed before the launch was queued. The arguments are now kept alive until the kernel has completed. That will matter once waiting for a PoCL kernel no longer blocks the garbage collector.

A second change to the back-end contract fixes a bug in how KernelInterface compiles closures. KI.@launch, and KA's generic launch, converted the callable with argconvert before passing it to kernel_function. The converted form of a closure only holds pointers to the arrays it captures, so nothing kept those arrays alive, and a kernel compiled with launch=false could read freed memory once the garbage collector had run:

function mk(backend, out)
    a = KI.allocate(backend, Int32, 1); fill!(a, 42)
    KI.@launch backend launch=false (() -> (@inbounds out[1] = a[1]; nothing))()
end
kernel = mk(backend, out)
GC.gc(true)
kernel()   # `a` may have been freed

A back end also couldn't see the captured arrays at launch, which Metal needs, to declare them to the command encoder (Metal.jl#992 fixed the same for @metal). kernel_function now receives the original callable and converts it itself, and the kernel it returns keeps the original alive. A back end can convert it again at every launch if it needs to. KernelInterface's testsuite checks this with a kernel whose closure captures an array.

Per launch on PoCL, with Julia 1.13 and POCL_MAX_PTHREAD_COUNT=1:

main this PR
4 arguments 32 B, 2 allocations 32 B, 2 allocations
40 arguments 25.7 KB, 148 allocations 32 B, 2 allocations

The launch testsuite that back ends run now includes a kernel with 40 arguments, and the PoCL tests check that launching it allocates no more than launching one with 4. The KernelInterface tests check that the arguments reach KI.launch as one tuple, including a single tuple-valued argument. Tested on Julia 1.10 and 1.13.

Each layer of PoCL's launch path took the kernel arguments as varargs and
splatted them into the next: the kernel object's call, `clcall`, `call`, and
`set_args!`, which recursed with one splat per argument. Julia doesn't turn a
splat of more than 32 elements into a direct call, and a method with both
varargs and keyword arguments splats them into its body, so launching a kernel
with many arguments went through `Core._apply_iterate` several times.

Pass the arguments along as one tuple instead, and generate the per-argument
code. The kernel object keeps its call syntax, with its keyword method defined
explicitly, and `launch_tuple` takes the tuple directly. `@opencl` builds the
argument types without `map`, which isn't type stable for 32 or more elements.
The back-end hook was `launch(kernel, groups, items, args...; kwargs...)`.
A method with both varargs and keyword arguments splats the arguments into its
body, and Julia doesn't turn a splat of more than 32 elements into a direct
call, so a launch with many arguments allocated in every back end, whatever the
caller did. Avoiding that meant writing a `Core.kwcall` method by hand in each
back end.

Make the hook take the arguments as one tuple instead:

    launch(kernel::Kernel{<:NewBackend}, groups::Dims{3}, items::Dims{3}, args::Tuple; kwargs...)

A back end can then pass the tuple on to its own launcher. This is breaking, so
this is KernelInterface 0.4. No back end depends on 0.3 yet.

Calling a `Kernel` keeps its syntax. Its keyword method is defined explicitly,
so that it can pass the arguments on as a tuple, and `@launch` builds the
argument types without `map`, which isn't type stable for 32 or more elements.
@maleadt

maleadt commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Could make this 0.3.1 since no back-end has been adapted yet, but it doesn't hurt being semver-strict.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main b2a4ee1... main / b2a4ee1...
const/@Const/Float32/262144 0.285 ± 0.026 ms 0.281 ± 0.027 ms 1.02 ± 0.13
const/@Const/Float32/65536 0.115 ± 0.024 ms 0.112 ± 0.024 ms 1.02 ± 0.3
const/@Const/Float64/262144 0.427 ± 0.028 ms 0.423 ± 0.03 ms 1.01 ± 0.098
const/@Const/Float64/65536 0.166 ± 0.013 ms 0.167 ± 0.014 ms 0.999 ± 0.12
const/unmarked/Float32/262144 1.46 ± 0.03 ms 1.14 ± 0.038 ms 1.28 ± 0.051
const/unmarked/Float32/65536 0.319 ± 0.025 ms 0.319 ± 0.031 ms 1 ± 0.13
const/unmarked/Float64/262144 1.36 ± 0.034 ms 1.36 ± 0.036 ms 1 ± 0.037
const/unmarked/Float64/65536 0.374 ± 0.038 ms 0.393 ± 0.04 ms 0.953 ± 0.14
launch/3D static workgroup, dynamic ndrange 0.0738 ± 0.015 ms 0.0749 ± 0.016 ms 0.986 ± 0.29
launch/3D static workgroup, static ndrange 0.0738 ± 0.013 ms 0.0744 ± 0.014 ms 0.992 ± 0.26
launch/dynamic workgroup, dynamic ndrange 0.072 ± 0.032 ms 0.0708 ± 0.031 ms 1.02 ± 0.63
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 0.0719 ± 0.03 ms 0.0686 ± 0.036 ms 1.05 ± 0.7
launch/static workgroup, dynamic ndrange 0.0739 ± 0.021 ms 0.0735 ± 0.018 ms 1.01 ± 0.38
launch/static workgroup, static ndrange 0.0733 ± 0.021 ms 0.0735 ± 0.017 ms 0.997 ± 0.36
partition/dynamic workgroup, dynamic ndrange 0.0592 ± 0.00084 μs 0.0589 ± 0.0022 μs 1.01 ± 0.04
partition/static workgroup, dynamic ndrange 0.0646 ± 0.012 μs 0.0665 ± 0.013 μs 0.972 ± 0.26
partition/static workgroup, static ndrange 2.47 ± 0.01 ns 1.55 ± 0.009 ns 1.59 ± 0.011
saxpy/default/Float16/1024 0.0748 ± 0.0045 ms 0.0747 ± 0.0088 ms 1 ± 0.13
saxpy/default/Float16/1048576 0.802 ± 0.041 ms 0.797 ± 0.047 ms 1.01 ± 0.078
saxpy/default/Float16/16384 0.0655 ± 0.034 ms 0.0619 ± 0.032 ms 1.06 ± 0.77
saxpy/default/Float16/2048 0.0762 ± 0.017 ms 0.0763 ± 0.017 ms 0.999 ± 0.32
saxpy/default/Float16/256 0.0748 ± 0.018 ms 0.0746 ± 0.028 ms 1 ± 0.45
saxpy/default/Float16/262144 0.253 ± 0.035 ms 0.251 ± 0.035 ms 1.01 ± 0.2
saxpy/default/Float16/32768 0.0748 ± 0.032 ms 0.0726 ± 0.032 ms 1.03 ± 0.63
saxpy/default/Float16/4096 0.0767 ± 0.031 ms 0.0745 ± 0.032 ms 1.03 ± 0.6
saxpy/default/Float16/512 0.075 ± 0.013 ms 0.0746 ± 0.023 ms 1.01 ± 0.36
saxpy/default/Float16/64 0.0753 ± 0.015 ms 0.0752 ± 0.016 ms 1 ± 0.3
saxpy/default/Float16/65536 0.0964 ± 0.03 ms 0.0934 ± 0.03 ms 1.03 ± 0.46
saxpy/default/Float32/1024 0.0738 ± 0.021 ms 0.0738 ± 0.03 ms 0.999 ± 0.5
saxpy/default/Float32/1048576 0.388 ± 0.037 ms 0.384 ± 0.037 ms 1.01 ± 0.14
saxpy/default/Float32/16384 0.0588 ± 0.032 ms 0.0552 ± 0.032 ms 1.07 ± 0.84
saxpy/default/Float32/2048 0.0742 ± 0.023 ms 0.0747 ± 0.013 ms 0.994 ± 0.35
saxpy/default/Float32/256 0.0744 ± 0.021 ms 0.0743 ± 0.032 ms 1 ± 0.52
saxpy/default/Float32/262144 0.144 ± 0.029 ms 0.142 ± 0.031 ms 1.01 ± 0.3
saxpy/default/Float32/32768 0.0647 ± 0.031 ms 0.0623 ± 0.031 ms 1.04 ± 0.72
saxpy/default/Float32/4096 0.0765 ± 0.017 ms 0.0758 ± 0.026 ms 1.01 ± 0.41
saxpy/default/Float32/512 0.0739 ± 0.018 ms 0.0739 ± 0.029 ms 0.999 ± 0.46
saxpy/default/Float32/64 0.0748 ± 0.021 ms 0.0746 ± 0.026 ms 1 ± 0.45
saxpy/default/Float32/65536 0.0766 ± 0.03 ms 0.0745 ± 0.031 ms 1.03 ± 0.59
saxpy/default/Float64/1024 0.0739 ± 0.02 ms 0.0734 ± 0.024 ms 1.01 ± 0.43
saxpy/default/Float64/1048576 0.659 ± 0.085 ms 0.606 ± 0.069 ms 1.09 ± 0.19
saxpy/default/Float64/16384 0.0571 ± 0.029 ms 0.0567 ± 0.031 ms 1.01 ± 0.75
saxpy/default/Float64/2048 0.0735 ± 0.028 ms 0.0699 ± 0.032 ms 1.05 ± 0.62
saxpy/default/Float64/256 0.0739 ± 0.019 ms 0.0742 ± 0.015 ms 0.997 ± 0.32
saxpy/default/Float64/262144 0.206 ± 0.035 ms 0.198 ± 0.037 ms 1.04 ± 0.26
saxpy/default/Float64/32768 0.0702 ± 0.031 ms 0.068 ± 0.031 ms 1.03 ± 0.65
saxpy/default/Float64/4096 0.0703 ± 0.027 ms 0.0682 ± 0.028 ms 1.03 ± 0.58
saxpy/default/Float64/512 0.074 ± 0.016 ms 0.0736 ± 0.022 ms 1 ± 0.37
saxpy/default/Float64/64 0.0752 ± 0.014 ms 0.0752 ± 0.023 ms 1 ± 0.36
saxpy/default/Float64/65536 0.0936 ± 0.032 ms 0.0934 ± 0.032 ms 1 ± 0.48
saxpy/static workgroup=(1024,)/Float16/1024 0.0736 ± 0.023 ms 0.0741 ± 0.026 ms 0.994 ± 0.47
saxpy/static workgroup=(1024,)/Float16/1048576 0.802 ± 0.043 ms 0.797 ± 0.04 ms 1.01 ± 0.074
saxpy/static workgroup=(1024,)/Float16/16384 0.0639 ± 0.033 ms 0.0614 ± 0.032 ms 1.04 ± 0.76
saxpy/static workgroup=(1024,)/Float16/2048 0.0755 ± 0.024 ms 0.0759 ± 0.022 ms 0.995 ± 0.42
saxpy/static workgroup=(1024,)/Float16/256 0.0732 ± 0.027 ms 0.0743 ± 0.025 ms 0.986 ± 0.49
saxpy/static workgroup=(1024,)/Float16/262144 0.247 ± 0.035 ms 0.246 ± 0.034 ms 1.01 ± 0.2
saxpy/static workgroup=(1024,)/Float16/32768 0.0744 ± 0.031 ms 0.0728 ± 0.031 ms 1.02 ± 0.61
saxpy/static workgroup=(1024,)/Float16/4096 0.0717 ± 0.032 ms 0.0689 ± 0.032 ms 1.04 ± 0.67
saxpy/static workgroup=(1024,)/Float16/512 0.0743 ± 0.018 ms 0.0744 ± 0.023 ms 0.998 ± 0.39
saxpy/static workgroup=(1024,)/Float16/64 0.0738 ± 0.024 ms 0.0744 ± 0.016 ms 0.992 ± 0.39
saxpy/static workgroup=(1024,)/Float16/65536 0.0973 ± 0.03 ms 0.0941 ± 0.028 ms 1.03 ± 0.44
saxpy/static workgroup=(1024,)/Float32/1024 0.0736 ± 0.013 ms 0.074 ± 0.02 ms 0.994 ± 0.32
saxpy/static workgroup=(1024,)/Float32/1048576 0.401 ± 0.037 ms 0.402 ± 0.033 ms 0.998 ± 0.12
saxpy/static workgroup=(1024,)/Float32/16384 0.0584 ± 0.03 ms 0.0573 ± 0.029 ms 1.02 ± 0.73
saxpy/static workgroup=(1024,)/Float32/2048 0.0746 ± 0.018 ms 0.0745 ± 0.022 ms 1 ± 0.38
saxpy/static workgroup=(1024,)/Float32/256 0.0741 ± 0.025 ms 0.0745 ± 0.026 ms 0.995 ± 0.49
saxpy/static workgroup=(1024,)/Float32/262144 0.151 ± 0.03 ms 0.15 ± 0.028 ms 1.01 ± 0.28
saxpy/static workgroup=(1024,)/Float32/32768 0.0655 ± 0.029 ms 0.0626 ± 0.028 ms 1.05 ± 0.66
saxpy/static workgroup=(1024,)/Float32/4096 0.0721 ± 0.028 ms 0.0754 ± 0.027 ms 0.957 ± 0.51
saxpy/static workgroup=(1024,)/Float32/512 0.0734 ± 0.018 ms 0.0743 ± 0.014 ms 0.987 ± 0.31
saxpy/static workgroup=(1024,)/Float32/64 0.0747 ± 0.011 ms 0.0746 ± 0.023 ms 1 ± 0.34
saxpy/static workgroup=(1024,)/Float32/65536 0.08 ± 0.03 ms 0.0764 ± 0.029 ms 1.05 ± 0.56
saxpy/static workgroup=(1024,)/Float64/1024 0.0741 ± 0.018 ms 0.0738 ± 0.022 ms 1 ± 0.39
saxpy/static workgroup=(1024,)/Float64/1048576 0.601 ± 0.073 ms 0.597 ± 0.067 ms 1.01 ± 0.17
saxpy/static workgroup=(1024,)/Float64/16384 0.0624 ± 0.028 ms 0.0611 ± 0.029 ms 1.02 ± 0.67
saxpy/static workgroup=(1024,)/Float64/2048 0.0722 ± 0.03 ms 0.0723 ± 0.03 ms 0.999 ± 0.58
saxpy/static workgroup=(1024,)/Float64/256 0.0737 ± 0.022 ms 0.0738 ± 0.026 ms 0.997 ± 0.46
saxpy/static workgroup=(1024,)/Float64/262144 0.204 ± 0.035 ms 0.204 ± 0.037 ms 1 ± 0.25
saxpy/static workgroup=(1024,)/Float64/32768 0.0752 ± 0.029 ms 0.0714 ± 0.029 ms 1.05 ± 0.59
saxpy/static workgroup=(1024,)/Float64/4096 0.0665 ± 0.029 ms 0.0614 ± 0.031 ms 1.08 ± 0.72
saxpy/static workgroup=(1024,)/Float64/512 0.0736 ± 0.018 ms 0.0737 ± 0.011 ms 0.999 ± 0.29
saxpy/static workgroup=(1024,)/Float64/64 0.074 ± 0.023 ms 0.0746 ± 0.021 ms 0.992 ± 0.42
saxpy/static workgroup=(1024,)/Float64/65536 0.0969 ± 0.032 ms 0.0929 ± 0.031 ms 1.04 ± 0.49
time_to_load 0.779 ± 0.0083 s 0.782 ± 0.014 s 0.996 ± 0.021
main b2a4ee1... main / b2a4ee1...
const/@Const/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float64/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
launch/3D static workgroup, dynamic ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/3D static workgroup, static ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/dynamic workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, static ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

The generic launch splatted the kernel arguments from the `Kernel` call down to
the `KI.Kernel` call, and mapped over them to compute the argument types. With
more than 32 arguments neither is a direct call, so on PoCL a kernel with 40
arguments allocated 25.7 KB per launch, against 32 bytes with 4 arguments.

Pass them along as a tuple, as `KI.Kernel` now does, and generate the code
that computes the argument types and calls the `KI.Kernel`. A 40-argument
launch now allocates as much as a 4-argument one.

CUDA.jl#3309 did the same for CUDA's own launch of KA kernels, which the
generic launch replaces.
Converting an `Array` argument for PoCL yields a device array that only holds a
pointer, and nothing kept the original arguments alive after they were
converted. The garbage collector could therefore free an array while its
remaining arguments were converted, or while the kernel ran. A launch through
KernelAbstractions happens to keep them rooted, but calling a PoCL kernel
directly doesn't: with a forced collection during argument conversion, the
array passed to the kernel was finalized before the launch was queued.

Preserve the arguments while converting them and queuing the launch, and in
`KI.launch`, which waits for the kernel, until it completes. The latter matters
once waiting no longer blocks the garbage collector.
`KI.@launch` and KernelAbstractions' generic launch converted the callable with
`argconvert` before passing it to `kernel_function`. For a closure, the
converted callable only holds pointers to the arrays it captures, so nothing
kept those arrays alive: a kernel compiled with `launch=false` could read freed
memory once the garbage collector had run. A back end also couldn't see the
captured arrays at launch, which Metal needs to declare them to the command
encoder, as `@metal` has done since Metal.jl#992.

`kernel_function` now receives the original callable and converts it itself,
and the kernel keeps the original alive. PoCL's kernels keep it next to the
compiled kernel, and root it while launching. KernelInterface's testsuite checks
that a kernel whose closure captures an array still works after a collection.
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.64286% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 69.54%. Comparing base (d78cad4) to head (b2a4ee1).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/pocl/compiler/execution.jl 83.33% 2 Missing ⚠️
src/pocl/nanoOpenCL.jl 93.33% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #811      +/-   ##
==========================================
+ Coverage   69.46%   69.54%   +0.07%     
==========================================
  Files          26       26              
  Lines        2142     2157      +15     
==========================================
+ Hits         1488     1500      +12     
- Misses        654      657       +3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant