Skip to content

KernelInterface: sub-group shuffles, votes and a constant width - #3331

Draft
vchuravy wants to merge 2 commits into
ka-0.10from
vc/ki-subgroup-ops
Draft

vchuravy wants to merge 2 commits into
ka-0.10from
vc/ki-subgroup-ops

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026

Copy link
Copy Markdown
Member

Implements the sub-group communication of KernelInterface's updated contract from JuliaGPU/KernelAbstractions.jl#831 in the CUDA back-end (CUDACore/src/CUDAKernels.jl).

What's implemented

  • Shuffles KI.shfl, KI.shfl_down, KI.shfl_up, KI.shfl_xor, all with the full warp mask, like the existing shfl_down override.
    • The overrides, and KI.supports_shuffle, now only match const ShuffleTypes = Union{Bool, Base.BitInteger, Base.IEEEFloat}. These are the primitive types CUDA.jl's warp shuffles handle, splitting the wider ones into 32-bit shuffles. Other isbits types, Complex included, go to KI's generic fallback, which shuffles them field by field. Before this PR, shfl_down and supports_shuffle were generic where {T} methods.
    • Lanes and offsets are truncated with % UInt32 rather than converted, so an out-of-range value returns an unspecified value instead of throwing an InexactError. shfl_sync takes a 1-based lane and subtracts 1, so KI.shfl wraps the lane into 1:32 first, which is the same wrapping PTX applies to the 0-based lane.
  • Votes KI.sub_group_any, KI.sub_group_all, KI.sub_group_ballot. Ballot widens the UInt32 from vote_ballot_sync to UInt64. Bit i - 1 is set for lane i, since laneid() is 1-based.
  • Constant width: the warp size is now the constant 32i32 (WARP_SIZE) instead of a read of %WARP_SZ. KI.get_max_sub_group_size(T) returns 32 % T, a compile-time constant of the generated code as the contract requires. The other sub-group queries (get_num_sub_groups, get_sub_group_id, the partial-warp size) use the same constant. Every NVIDIA GPU has a warp size of 32, and CUDA.jl's warp intrinsics already hard-code it (ws = Int32(32) in warp.jl). The host-side KI.sub_group_size(::CUDABackend) still queries warpsize(device()), which returns 32.
KernelInterface CUDA.jl
shfl(val, lane) shfl_sync(FULL_MASK, val, lane′) with lane′ = ((lane - 1) % UInt32 & 0x1f) + 1
shfl_down(val, offset) shfl_down_sync(FULL_MASK, val, offset % UInt32)
shfl_up(val, offset) shfl_up_sync(FULL_MASK, val, offset % UInt32)
shfl_xor(val, mask) shfl_xor_sync(FULL_MASK, val, mask % UInt32)
sub_group_any(pred) vote_any_sync(FULL_MASK, pred)
sub_group_all(pred) vote_all_sync(FULL_MASK, pred)
sub_group_ballot(pred) UInt64(vote_ballot_sync(FULL_MASK, pred))
get_max_sub_group_size(T) 32i32 % T
supports_shuffle(::CUDABackend, ::Type{<:ShuffleTypes}) true (composites go to KI's fallback)

[TEMP] commit

The second commit, [TEMP] Get KernelAbstractions and KernelInterface from the vc/ki-subgroup-ops branch, changes the [sources] revs in CUDACore, CUDATools, lib/cusparse and test, and the Buildkite clone for Julia 1.10/1.11, from KA's main to vc/ki-subgroup-ops. This makes CI test against #831. Drop it, or squash it into the existing [TEMP] commit, once #831 is merged.

The CUDA back-end doesn't need the SPIRVIntrinsics source that KA's branch uses for POCL. Without it, the test environment resolves the registered SPIRVIntrinsics v1.2.0, and KA precompiles and loads fine.

Heads-up, a problem ka-0.10 already has: KA's main and #831 both require LLVM = "10", but ka-0.10 still has LLVM = "9.6" in CUDACore and "9.3.1" in CUDATools. main has since moved to LLVM.jl 10 (#3323). So ka-0.10 doesn't resolve against either KA branch, and CI on this PR will fail to resolve until ka-0.10 is rebased onto main. This PR doesn't fix that.

Local testing

I tested on a Quadro RTX 4000 (sm_75) with Julia 1.12.7. I used a throwaway local branch where ka-0.10, plus these commits, was rebased onto main; the only conflicts were the LLVM and version compat entries. It resolved KA/KI 0.10.0-dev/0.4.0-dev from #vc/ki-subgroup-ops (7f09a6c), LLVM v10.0.0 and GPUCompiler v2.11.1.

julia --project=test test/runtests.jl core/kernelabstractions core/kernelinterface

Result: Overall | 3612 pass, 17 broken, 3629 total, SUCCESS. The 17 broken tests are existing @test_brokens. This includes KI's testsuite from #831: the new shuffles (including structs), the votes and the constant width.

I also ran these checks by hand:

  • KI.shfl with lanes 0, -5, 33, typemax(Int), shfl_down(Int8, typemax(Int)) and shfl_up(UInt16, 40) don't throw.
  • The PTX for KI.get_max_sub_group_size() doesn't read %WARP_SZ.
  • KI.shfl on ComplexF64 goes through the KI fallback and gives the correct result.
  • supports_shuffle returns true for ComplexF64 and false for Char.

🤖 Generated with Claude Code

Implement the sub-group communication of KernelInterface's updated
contract: `shfl`, `shfl_up` and `shfl_xor` next to `shfl_down`, and the
votes `sub_group_any`, `sub_group_all` and `sub_group_ballot`, using the
warp intrinsics with the full mask.

The shuffles, and `supports_shuffle`, are restricted to the primitive
types CUDA's warp shuffles support, so that KernelInterface shuffles
structs and tuples (including `Complex`) field by field. Lanes and
offsets are truncated rather than converted, so that out-of-range values
give an unspecified value instead of throwing.

The warp size is now the constant 32 rather than a read of `%WARP_SZ`,
so that `get_max_sub_group_size` is a compile-time constant of the
generated code, as KernelInterface requires.

Assisted-by: Claude Code (Opus 5.5)
…roup-ops branch

Test against JuliaGPU/KernelAbstractions.jl#831. Drop this commit (going
back to KernelAbstractions' main branch) once that is merged.

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant