Launch @kernel kernels on any KernelInterface back end - #801
Merged
Merged
Conversation
maleadt
added this pull request to stack #802
September 28, 2026 09:19
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 09:40
9eceba1 to
ab061d4
Compare
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 10:09
ab061d4 to
c60e4ce
Compare
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 10:14
c60e4ce to
44f6987
Compare
maleadt
removed this pull request from stack #802
September 28, 2026 10:14
maleadt
added this pull request to stack #804
September 28, 2026 10:15
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 10:25
44f6987 to
a12f889
Compare
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 10:31
a12f889 to
6ec971e
Compare
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 28, 2026 11:00
bd6ef38 to
c178512
Compare
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 29, 2026 20:21
c178512 to
b9804b6
Compare
maleadt
removed this pull request from stack #804
September 29, 2026 20:22
maleadt
added this pull request to stack #808
September 29, 2026 20:22
maleadt
force-pushed
the
tb/ka-generic-launch
branch
2 times, most recently
from
September 30, 2026 04:57
c844b4d to
d34353a
Compare
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 30, 2026 08:16
d34353a to
172dbdd
Compare
This was referenced Sep 30, 2026
These couldn't be launched at all: `launch_config` dropped the static ndrange before partitioning with it as the preliminary workgroup size, and `partition` made the number of blocks static, so the context type changed when the workgroup size was tuned after compiling the kernel. The number of blocks is now only static if the workgroup size is too. `launch_config` also partitions with the given ndrange, so that it is checked against the static one.
`(::KA.Kernel{<:KI.Backend})(args...; ndrange, workgroupsize)` now partitions
the ndrange, selects the launch, compiles the kernel with `KI.kernel_function`,
tunes the workgroup size with `KI.launch_configuration` (passing the number of
work-items as `nitems`) and launches it through the `KI.Kernel`, which
validates the sizes. Backends no longer copy `mkcontext`, `launch_config` and
the launch itself; they customize it with `KI.launch_configuration` and the
new `compiler_options` hook, and still implement `Scratchpad` and their Adapt
rules. A backend's own `(::KA.Kernel{MyBackend})` method still takes
precedence. PoCL uses the generic launch.
maleadt
force-pushed
the
tb/ka-generic-launch
branch
from
September 30, 2026 08:27
172dbdd to
063ab6a
Compare
This was referenced Sep 30, 2026
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/ndlaunch #801 +/- ##
===============================================
- Coverage 69.63% 69.46% -0.17%
===============================================
Files 25 26 +1
Lines 2134 2142 +8
===============================================
+ Hits 1486 1488 +2
- Misses 648 654 +6 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every backend currently launches
@kernelkernels with its own copy of the same ~100 lines: partitioning thendrange(launch_config), building the hidden context (mkcontext), compiling, tuning the workgroup size, and launching. CUDA.jl, Metal.jl, oneAPI.jl, AMDGPU.jl, OpenCL.jl and PoCL each have one. The copies have drifted: CUDA and oneAPI have aprefer_blocksoption, oneAPI a hand-written sizing heuristic, AMDGPU a different preliminary workgroup size. And #797 would have every one of them adopt N-d launches and 32-bit indexing separately.With KernelInterface 0.3 (#800), KI can compile, size and launch a kernel for any backend, so KA can do the rest. This PR moves the launch into KA (
src/backend_launch.jl), and PoCL is the first to use it. For a kernel on a KI backend, KA now:ndrange, and returns right away if it's empty;select_launch(Launch @kernel kernels on N-d grids, indexing in 32 bits #797), and compiles it withKI.kernel_functionfor a context that carries that choice;KI.launch_configurationfor one, passing the number of work-items. The launch doesn't depend on the tuned size, so the context type stays the same and the kernel isn't recompiled;KI.Kernel, which validates the sizes against the kernel's limits.A backend still provides, on top of KernelInterface:
KA.Scratchpadfor@private, e.g. anMArrayon CUDA or a stack allocation on PoCL;@Const;KI.launch_configuration. That's where CUDA's and oneAPI'sprefer_blocksgo, now that it receives the number of work-items asnitems;KA.compiler_options(kernel). CUDA uses it to passmaxthreadsfor a static workgroup size:To use the generic launch, a backend needs KernelInterface's typed index queries and an N-d
KI.launch, and it has to drop its overrides of KA's index functions. A backend that keeps its own(::KA.Kernel{MyBackend})method still takes precedence, so each port can switch on its own schedule.implementations.mdnow lists these hooks instead of #797's launch template, and describes overriding the launch as an escape hatch that relies on KA internals (select_launch,launch_workgroupsize).The first commit fixes a bug that every copy of the launch shared: kernels with a static
ndrangeand a tuned workgroup size couldn't be launched at all.launch_configdropped the staticndrangebefore using it as the preliminary workgroup size. Fixing only that isn't enough, sincepartitionthen made the number of blocks part of the context type, which changed when the workgroup size was tuned after compiling (aMethodErrorconverting the context). The number of blocks is now only static if the workgroup size is too.It also makes
launch_configcheck a givenndrangeagainst a static one, which it didn't do.