Conversation
- Add `sub_group_any`, `sub_group_all` (`cl_khr_subgroups`) and `sub_group_ballot` (`cl_khr_subgroup_ballot`). - Pass the lane of `sub_group_shuffle` and the mask of `sub_group_shuffle_xor` modulo `UInt32`: an out-of-range lane now gives an undefined value, as with the OpenCL C built-ins, instead of throwing an `InexactError`. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
This was referenced Oct 4, 2026
Add the collectives of `cl_khr_subgroups`: `sub_group_reduce_*`, `sub_group_scan_inclusive_*` and `sub_group_scan_exclusive_*` for `add`, `min` and `max`, and `sub_group_broadcast` (with a 1-based lane, like `sub_group_shuffle`), for 32- and 64-bit integers and floats. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
PoCL (7.2) loses the values of the work-items that computed them in the branch; mark that as broken there. Assisted-by: Claude Code (Opus 5.5)
The collectives (reductions, scans, broadcast, votes, ballot and shuffles) were declared through `@builtin_ccall` without the `convergent` attribute, so LLVM was free to make them control-dependent on additional values. For `r = sub_group_reduce_add(i <= m ? x[i] : 0)` the optimizer sank the call into both arms of the branch, so the work-items of a sub-group executed two different calls, which is undefined behavior. PoCL, which implements every collective with work-group barriers and scratch memory, then never ran the store of the work-items that took the first arm. Add a `convergent = true` option to `@builtin_ccall`, which declares the builtin `convergent` like the barriers, and use it for all collectives. The bounds-checked variant of the test remains broken on PoCL 7.2, where the (never taken) early exit of the bounds check makes WorkitemLoops peel the first work-item into a separate copy of the collective's scratch memory; that is fixed in PoCL by pocl/pocl@8fa5d2732. Assisted-by: Claude Code (Opus 5.5)
vchuravy
added a commit
to JuliaGPU/KernelAbstractions.jl
that referenced
this pull request
Oct 4, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes: SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2 gives a peeled work-item its own copy of a collective's scratch memory after a branch with an early exit, as bounds checks emit (fixed on PoCL's main branch, backport to 7.2 in pocl/pocl#2373). With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES` until `pocl_standalone_jll` includes the PoCL fix. Assisted-by: Claude Code (Opus 5.5)
This was referenced Oct 4, 2026
Without `convergent`, the optimizer duplicates `sub_group_any(a && b)` into both arms of the short-circuit (jump threading, when the predicate is used again after the vote), so the work-items of a sub-group call the vote separately. On PoCL every work-item then sees `false`. This was the cause of a silent NaN energy in Molly's tiled energy kernel on the PoCL back-end. The test fails on the parent of 0078f3a and passes with it. Assisted-by: Claude Code (Opus 5.5)
spirv2clc, which translates SPIR-V to OpenCL C for the OpenCL C program backend, doesn't implement OpGroupAll/Any/Broadcast/IAdd etc. or the GroupNonUniformBallot capability, so skip those tests on that backend. With pocl_jll 7.2.1+1, which includes pocl/pocl#2239, the bounds-checked divergent reduction passes on PoCL, but still fails intermittently as part of the test suite; skip it on PoCL for now. Assisted-by: Claude Code (Opus 5.5)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstreams two workarounds from KernelAbstractions' POCL back-end (JuliaGPU/KernelAbstractions.jl#831, see this review thread), where they implement KernelInterface's sub-group operations.
Votes
New
sub_group_any(::Bool)andsub_group_all(::Bool)(cl_khr_subgroups), andsub_group_ballot(::Bool)::NTuple{4, VecElement{UInt32}}(cl_khr_subgroup_ballot, OpenCL'suint4mask).These call the OpenCL C built-ins, which the SPIR-V back-end lowers, rather than
__spirv_GroupNonUniformAny/All/Ballot:@builtin_ccallcan't mangle theirboolarguments. Defining them here also registers the names inknown_intrinsicsproperly. From another package (KA's POCL back-end),@builtin_ccallregisters them at macro-expansion time into this package's array, and that registration doesn't survive precompilation.Shuffle lanes
sub_group_shuffleconverted the lane withUInt32(i - 1), which throws anInexactError(and leaves an error branch in the kernel) for an out-of-range lane, e.g. when implementing a shuffle-up assub_group_shuffle(x, lane - offset). The lane is now passed moduloUInt32, and so is the mask ofsub_group_shuffle_xor. An out-of-range lane gives an undefined value, as with the OpenCL C built-ins.Collectives
The
cl_khr_subgroupscollectives, which the SPIR-V back-end lowers toOpGroup*instructions, forInt32,UInt32,Int64,UInt64,Float16,Float32andFloat64:sub_group_reduce_{add,min,max}sub_group_scan_inclusive_{add,min,max}sub_group_scan_exclusive_{add,min,max}sub_group_broadcast(x, lane), with a 1-based lane likesub_group_shuffleKernelAbstractions uses these to implement
KernelInterface.sub_group_reduce/sub_group_scannatively.The non-uniform arithmetic of
cl_khr_subgroup_non_uniform_arithmetic(mul,and,or,xor, …) and the char/short types ofcl_khr_subgroup_extended_typesaren't included, because PoCL doesn't support them, so they couldn't be tested here.Convergence
All collectives (shuffles, votes, ballot, reductions, scans, broadcast) are now declared
convergent.@builtin_ccallgets aconvergent = trueoption for this. Without it, LLVM was free to make the calls control-dependent on more values. Forr = sub_group_reduce_add(i <= m ? x[i] : 0), the optimizer sank the call into both arms of the branch, so the work-items of a sub-group executed two different calls, which is undefined behavior. On PoCL, which implements the collectives with work-group barriers, this returned 0.A bounds-checked variant still fails on PoCL 7.2. There, the (never taken) early exit of the bounds check makes WorkitemLoops give the peeled first work-item its own scratch memory. PoCL fixed this on main in pocl/pocl@8fa5d2732 (#2239); Julia's PoCL builds carry it as a patch with JuliaPackaging/Yggdrasil#15001. The test marks that variant as broken on PoCL 7.2.
Tests
test/intrinsics.jl:any/allfor several predicate patternsballot, when the device hascl_khr_subgroup_ballotpocl/intrinsicsandintel/intrinsicspass (2238 tests). Both devices support sub-groups and ballot.KernelAbstractions can drop its copies once this is released, so a version bump of SPIRVIntrinsics (1.3.0, new features) would be needed.
🤖 Generated with Claude Code