KernelInterface: don't promise how work-items form sub-groups - #815
Merged
Merged
Conversation
KernelInterface promised that a work-group is divided into sub-groups of the sub-group width with at most one partial sub-group, i.e. `cld(prod(get_local_size()), get_max_sub_group_size())` of them. Only CUDA, HSA (AMD) and Metal specify linear packing of multi-dimensional work-groups. OpenCL requires that only the highest-numbered sub-group is smaller, but leaves the mapping implementation-defined, and SYCL, Vulkan and WebGPU don't specify either. Implementations differ: Intel's CPU OpenCL runtime forms sub-groups per row of a work-group by default, so a 33x2 work-group has four sub-groups of 32 and 1 work-items, and Intel's GPU runtime picks a hardware walk order that isn't x-fastest for some shapes. Require only what backends can ensure everywhere: unique and invariant (sub-group, lane) pairs, dense ids and lanes, and that a 1-D work-group of at most the sub-group width is a single sub-group. `get_num_sub_groups` is at least the count full sub-groups would need, but can be larger, so storage for a value per sub-group has to allow for one per work-item. The testsuite checks those invariants over 1-D, 2-D and 3-D work-groups whose size is or isn't a multiple of the width, and combines a value per sub-group through local memory, as reductions do. The shuffle test no longer assumes that a work-group of twice the width consists of two full sub-groups.
maleadt
added a commit
to JuliaGPU/OpenCL.jl
that referenced
this pull request
Sep 30, 2026
Implement the sub-group queries, `sub_group_barrier` and `shfl_down` with OpenCL's sub-group builtins, whose semantics already match KernelInterface's: the last sub-group of a work-group can be partial, and `get_sub_group_size` counts the work-items that are present. KernelInterface requires kernels to execute with the sub-group width that `KI.sub_group_size` reports. `clfunction` requests that width on devices with `cl_intel_required_subgroup_size`, so only those report sub-group support, and shuffles for the types that `cl_khr_subgroup_shuffle` covers on the device. A `sub_group_size` compiler option asking for another width is rejected. How work-items form sub-groups differs between devices: Intel's CPU runtime forms them per row of a multi-dimensional work-group. KernelInterface doesn't promise more than that (JuliaGPU/KernelAbstractions.jl#815). Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #815 +/- ##
=======================================
Coverage 69.54% 69.54%
=======================================
Files 26 26
Lines 2157 2157
=======================================
Hits 1500 1500
Misses 657 657 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
KernelInterface promised that a work-group is divided into sub-groups of the sub-group width, the last of which can be partial, so that there are
cld(prod(get_local_size()), get_max_sub_group_size())of them. Porting OpenCL.jl (JuliaGPU/OpenCL.jl#519) showed that this doesn't hold everywhere: Intel's CPU OpenCL runtime forms sub-groups per row of a multi-dimensional work-group, so a 33×2 work-group consists of four sub-groups of 32 and 1 work-items, not three.Recording
get_sub_group_id()andget_sub_group_size()for every work-item of a 33×2 work-group, with OpenCL.jl and sub-groups of 32:That isn't a quirk of one runtime. Only CUDA (Programming Guide §3.2.2.2), HSA (PRM §2.6, which AMD's hardware follows) and Metal (MSL §5.2.3.6) guarantee linear, x-fastest packing for multi-dimensional work-groups. The others leave it implementation-defined:
GetGroupWaveCountis "at least ceil(...), but may be larger".Implementations use that freedom. Besides Intel's CPU runtime, which has a setting for it (
CL_CONFIG_CPU_SUB_GROUP_CONSTRUCTION, per row by default), Intel's GPU runtime chooses a hardware walk order that isn't x-fastest for some work-group shapes on recent GPUs (Arc, Ponte Vecchio, Meteor Lake and later), and Mesa's Intel and Adreno drivers tile some shaders. Software that relies on linear packing in multi-dimensional work-groups (CUB, rocPRIM, Kokkos' CUDA back end) only targets CUDA and HIP; portable libraries use 1-D work-groups, or only a sub-group's own ids. Nothing in the Julia ecosystem uses KernelInterface's sub-groups yet, so this is the time to fix the contract.This PR makes KernelInterface promise only what every backend can ensure. Backends that report sub-group support have to guarantee that:
(get_sub_group_id(), get_sub_group_local_id())pair in its work-group, which doesn't change during the kernel;1:get_num_sub_groups(), and the lanes of a sub-group are1:get_sub_group_size();sub_group_size(backend)work-items is a single sub-group, which is what kernels that use one sub-group per work-group need.Which work-items form a sub-group, and which sub-groups are partial, is unspecified.
get_num_sub_groups()is at least the count that full sub-groups would need, but can be larger. The price is that storage for a value per sub-group, e.g. the partial results of a reduction, has to allow for one per work-item, and the code combining them has to useget_num_sub_groups()rather than compute the count. That's what the new test does:The fixed width that
sub_group_size(backend)promises is unchanged: kernels still execute with that maximum width. Splitting support for sub-group operations from that guarantee (#810, item 4) is left for another PR.The testsuite now checks those guarantees over 1-D work-groups around the width, and over 2-D and 3-D ones whose first dimension is or isn't a multiple of the width, and combines a value per sub-group as above. The shuffle test no longer assumes that a work-group of twice the width consists of two full sub-groups. It all passes on PoCL, on CUDA, and on Intel's CPU OpenCL runtime with sub-group support reported. After this, OpenCL.jl#519 no longer has to hide sub-groups on that runtime.