Launch @kernel kernels on N-d grids, indexing in 32 bits - #797
Merged
Merged
Conversation
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 27, 2026 12:28
f51c0d9 to
c006f4e
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
4 times, most recently
from
September 27, 2026 13:59
a9b94b9 to
523886d
Compare
vchuravy
approved these changes
Sep 27, 2026
This comment was marked as resolved.
This comment was marked as resolved.
maleadt
force-pushed
the
tb/ndlaunch
branch
2 times, most recently
from
September 27, 2026 20:12
ff75616 to
b7cf4cb
Compare
maleadt
added this pull request to stack #799
September 27, 2026 20:17
This comment was marked as resolved.
This comment was marked as resolved.
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 28, 2026 05:09
b7cf4cb to
e7b8fdf
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
2 times, most recently
from
September 28, 2026 09:18
f1cccb3 to
5e0f16e
Compare
maleadt
removed this pull request from stack #799
September 28, 2026 09:19
maleadt
added this pull request to stack #802
September 28, 2026 09:19
maleadt
force-pushed
the
tb/ndlaunch
branch
2 times, most recently
from
September 28, 2026 10:09
5b9fb86 to
bdf401d
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 28, 2026 10:14
bdf401d to
1b18464
Compare
maleadt
added this pull request to stack #804
September 28, 2026 10:15
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 28, 2026 10:25
1b18464 to
f7d94a7
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 28, 2026 10:31
f7d94a7 to
5e13c50
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 28, 2026 11:00
5e13c50 to
5a583d0
Compare
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 29, 2026 20:21
5a583d0 to
d4833ae
Compare
maleadt
removed this pull request from stack #804
September 29, 2026 20:22
maleadt
added this pull request to stack #808
September 29, 2026 20:22
maleadt
force-pushed
the
tb/ndlaunch
branch
3 times, most recently
from
September 30, 2026 08:16
a5a53b9 to
8f123c8
Compare
`@index` assumes that a kernel was launched on a 1-D grid, so every work-item
decomposes its linear group and local id into Cartesian positions, with 64-bit
integer divisions for a dynamically-sized ndrange.
Backends can now record how they launched a kernel in its context, and `@index`
specializes on that:
- `NDLaunch{T}`: the grid has the shape of the iteration space (for as many
dimensions as the backend's grid has), so the hardware ids are the Cartesian
positions and no divisions are needed;
- `LinearLaunch{T}`: the 1-D grid used so far.
Either way `@index` computes in `T`, `Int32` whenever the padded iteration
space fits, and still returns `Int`s.
`select_launch` chooses the launch from the iteration space and the backend's
limits. When the workgroup size will be tuned, the choice holds for every
workgroup size tuning can pick, so the context type (and thus the compiled
kernel) doesn't change. Contexts without a launch keep the existing behavior,
so backends opt in; `__validindex` is now generic, dispatching on the launch.
PoCL launches on N-d grids.
The default path decomposed the hardware ids by indexing `blocks(iterspace)` and
`workitems(iterspace)`, which throws for 0-d ndranges (`linear_index` reads
`I.I[0]`), and which under `--check-bounds=yes` puts bounds checks depending on
the local id in front of `@synchronize`. POCL crashes on the resulting barrier
in non-uniform control flow, e.g. for a 3-D ndrange with partial workgroups.
Computing the positions like a `LinearLaunch{Int}` avoids both, and leaves
only one implementation of the index functions.
maleadt
force-pushed
the
tb/ndlaunch
branch
from
September 30, 2026 08:27
8f123c8 to
7509cf4
Compare
This was referenced Sep 30, 2026
Intel's CPU runtime skips the partial last work-group along x in N-d launches
JuliaGPU/OpenCL.jl#521
Open
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@indexassumes that a kernel was launched on a 1-D grid. To get Cartesian indices, every work-item decomposes its linear group and local id with 64-bit integer divisions: 2(N−1) of them for a dynamically-sized N-dndrange, and the same again for the bounds check (#470, #396). On GPUs, that is often most of a kernel's index arithmetic.This PR lets a backend launch a kernel on a grid with the shape of the iteration space, and lets
@indexcompute in 32 bits. With the CUDA.jl side of it (JuliaGPU/CUDA.jl#3304), on an RTX 5080, compared to the KA 0.10 port (JuliaGPU/CUDA.jl#3302), both with the inlining fix from #798:A .* v(256³)A .+ vA .* vpermutedims3-D1-D and memory-bound 2-D kernels don't change. Oceananigans runs 1–3% faster per F32 time step. There are more numbers in the CUDA.jl PR.
How it works
The hidden context argument of a
@kernel(CompilerMetadata) gets alaunchfield that says how the backend launched the kernel, and the index functions specialize on it:NDLaunch{T}: the grid has the shape of the iteration space, so the hardware group and local ids are the Cartesian positions and no divisions are needed. This works for as many dimensions as the hardware grid has (3, according to the KernelInterface limits); larger iteration spaces use a linear launch.LinearLaunch{T}: the 1-D grid used so far.Either way,
@indexcomputes inT:Int32whenever the padded iteration space fits,Intotherwise. It still returnsInts. Linear indices keep their x-fastest order, so sub-groups still hold consecutive work-items.KA.select_launchchooses the launch from the iteration space and the backend's limits (KI.max_work_group_size,KI.max_work_group_dimsandKI.max_num_groups). When the workgroup size is tuned after compiling, the choice has to hold for every size that tuning can pick, because the launch is part of the context type and thus of the compiled kernel. It does, because tuning distributes work-items the same wayselect_launchassumes (KA.launch_workgroupsize). An iteration space with more thantypemax(Int)work-items is now rejected with anArgumentError.Adopting this is optional for a backend: a context without a launch is indexed as before (now with the same code, as a
LinearLaunch{Int}). PoCL adopts it here.docs/src/implementations.mddescribes what a backend's launch has to do; #801 then moves that launch into KA, so that backends don't have to.Packages that customize the iteration space (a custom
partitionorexpand, like Oceananigans) keep working: only the iteration spaces KA creates itself take the direct path, and the launch is chosen from the iteration space that is actually launched. AnAdaptrule forCompilerMetadatahas to pass thelaunchalong, though.The second commit also fixes two bugs in the default indexing path: 0-d ndranges threw a bounds error, and under
--check-bounds=yesthe index bounds checks crashed PoCL's compiler for kernels with@synchronize(pocl/pocl#2345).