Skip to content

SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes - #526

Open
vchuravy wants to merge 6 commits into
mainfrom
vc/subgroup-votes
Open

vchuravy wants to merge 6 commits into
mainfrom
vc/subgroup-votes

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Upstreams two workarounds from KernelAbstractions' POCL back-end (JuliaGPU/KernelAbstractions.jl#831, see this review thread), where they implement KernelInterface's sub-group operations.

Votes

New sub_group_any(::Bool) and sub_group_all(::Bool) (cl_khr_subgroups), and sub_group_ballot(::Bool)::NTuple{4, VecElement{UInt32}} (cl_khr_subgroup_ballot, OpenCL's uint4 mask).

These call the OpenCL C built-ins, which the SPIR-V back-end lowers, rather than __spirv_GroupNonUniformAny/All/Ballot: @builtin_ccall can't mangle their bool arguments. Defining them here also registers the names in known_intrinsics properly. From another package (KA's POCL back-end), @builtin_ccall registers them at macro-expansion time into this package's array, and that registration doesn't survive precompilation.

Shuffle lanes

sub_group_shuffle converted the lane with UInt32(i - 1), which throws an InexactError (and leaves an error branch in the kernel) for an out-of-range lane, e.g. when implementing a shuffle-up as sub_group_shuffle(x, lane - offset). The lane is now passed modulo UInt32, and so is the mask of sub_group_shuffle_xor. An out-of-range lane gives an undefined value, as with the OpenCL C built-ins.

Collectives

The cl_khr_subgroups collectives, which the SPIR-V back-end lowers to OpGroup* instructions, for Int32, UInt32, Int64, UInt64, Float16, Float32 and Float64:

  • sub_group_reduce_{add,min,max}
  • sub_group_scan_inclusive_{add,min,max}
  • sub_group_scan_exclusive_{add,min,max}
  • sub_group_broadcast(x, lane), with a 1-based lane like sub_group_shuffle

KernelAbstractions uses these to implement KernelInterface.sub_group_reduce/sub_group_scan natively.

The non-uniform arithmetic of cl_khr_subgroup_non_uniform_arithmetic (mul, and, or, xor, …) and the char/short types of cl_khr_subgroup_extended_types aren't included, because PoCL doesn't support them, so they couldn't be tested here.

Convergence

All collectives (shuffles, votes, ballot, reductions, scans, broadcast) are now declared convergent. @builtin_ccall gets a convergent = true option for this. Without it, LLVM was free to make the calls control-dependent on more values. For r = sub_group_reduce_add(i <= m ? x[i] : 0), the optimizer sank the call into both arms of the branch, so the work-items of a sub-group executed two different calls, which is undefined behavior. On PoCL, which implements the collectives with work-group barriers, this returned 0.

A bounds-checked variant still fails on PoCL 7.2. There, the (never taken) early exit of the bounds check makes WorkitemLoops give the peeled first work-item its own scratch memory. PoCL fixed this on main in pocl/pocl@8fa5d2732 (#2239); Julia's PoCL builds carry it as a patch with JuliaPackaging/Yggdrasil#15001. The test marks that variant as broken on PoCL 7.2.

Tests

  • New tests in test/intrinsics.jl:
    • a shuffle from an out-of-range lane
    • any/all for several predicate patterns
    • ballot, when the device has cl_khr_subgroup_ballot
    • the reductions, scans and broadcast
  • Local results: pocl/intrinsics and intel/intrinsics pass (2238 tests). Both devices support sub-groups and ballot.

KernelAbstractions can drop its copies once this is released, so a version bump of SPIRVIntrinsics (1.3.0, new features) would be needed.

🤖 Generated with Claude Code

- Add `sub_group_any`, `sub_group_all` (`cl_khr_subgroups`) and
  `sub_group_ballot` (`cl_khr_subgroup_ballot`).
- Pass the lane of `sub_group_shuffle` and the mask of `sub_group_shuffle_xor`
  modulo `UInt32`: an out-of-range lane now gives an undefined value, as with
  the OpenCL C built-ins, instead of throwing an `InexactError`.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
Replace the local workarounds with the votes and unchecked shuffle lanes from
JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through
`[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the
Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's
ka-0.10 branch).

Assisted-by: Claude Code (Opus 5.5)
Add the collectives of `cl_khr_subgroups`: `sub_group_reduce_*`,
`sub_group_scan_inclusive_*` and `sub_group_scan_exclusive_*` for `add`, `min`
and `max`, and `sub_group_broadcast` (with a 1-based lane, like
`sub_group_shuffle`), for 32- and 64-bit integers and floats.

Assisted-by: Claude Code (Opus 5.5)
@vchuravy vchuravy changed the title SPIRVIntrinsics: sub-group votes, unchecked shuffle lanes SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes Oct 4, 2026
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of
`cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on
32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the
fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero
differently.

Test the operators and types that backends may implement natively, including a
NaN, and that POCL uses the native reduction.

Assisted-by: Claude Code (Opus 5.5)
PoCL (7.2) loses the values of the work-items that computed them in the
branch; mark that as broken there.

Assisted-by: Claude Code (Opus 5.5)
The collectives (reductions, scans, broadcast, votes, ballot and shuffles)
were declared through `@builtin_ccall` without the `convergent` attribute,
so LLVM was free to make them control-dependent on additional values. For
`r = sub_group_reduce_add(i <= m ? x[i] : 0)` the optimizer sank the call into
both arms of the branch, so the work-items of a sub-group executed two
different calls, which is undefined behavior. PoCL, which implements every
collective with work-group barriers and scratch memory, then never ran the
store of the work-items that took the first arm.

Add a `convergent = true` option to `@builtin_ccall`, which declares the
builtin `convergent` like the barriers, and use it for all collectives.

The bounds-checked variant of the test remains broken on PoCL 7.2, where the
(never taken) early exit of the bounds check makes WorkitemLoops peel the
first work-item into a separate copy of the collective's scratch memory;
that is fixed in PoCL by pocl/pocl@8fa5d2732.

Assisted-by: Claude Code (Opus 5.5)
vchuravy added a commit to JuliaGPU/KernelAbstractions.jl that referenced this pull request Oct 4, 2026
The wrong results of the native `cl_khr_subgroups` collectives had two causes:
SPIRVIntrinsics declared them without `convergent`, so LLVM duplicated the
calls into divergent branches (fixed in JuliaGPU/OpenCL.jl#526), and PoCL 7.2
gives a peeled work-item its own copy of a collective's scratch memory after a
branch with an early exit, as bounds checks emit (fixed on PoCL's main branch,
backport to 7.2 in pocl/pocl#2373).

With both fixed, i.e. with `POCL_WORK_GROUP_METHOD=cbs` for now, the native
reductions and scans pass the tests. Keep them behind `NATIVE_COLLECTIVES`
until `pocl_standalone_jll` includes the PoCL fix.

Assisted-by: Claude Code (Opus 5.5)
Without `convergent`, the optimizer duplicates `sub_group_any(a && b)` into
both arms of the short-circuit (jump threading, when the predicate is used
again after the vote), so the work-items of a sub-group call the vote
separately. On PoCL every work-item then sees `false`. This was the cause of a
silent NaN energy in Molly's tiled energy kernel on the PoCL back-end.

The test fails on the parent of 0078f3a and passes with it.

Assisted-by: Claude Code (Opus 5.5)
spirv2clc, which translates SPIR-V to OpenCL C for the OpenCL C program
backend, doesn't implement OpGroupAll/Any/Broadcast/IAdd etc. or the
GroupNonUniformBallot capability, so skip those tests on that backend. With
pocl_jll 7.2.1+1, which includes pocl/pocl#2239, the bounds-checked divergent
reduction passes on PoCL, but still fails intermittently as part of the test
suite; skip it on PoCL for now.

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant