Conversation
Implement the shuffles `shfl`, `shfl_down`, `shfl_up` and `shfl_xor` and the votes `sub_group_any`, `sub_group_all` and `sub_group_ballot` of KernelInterface. The shuffles are `ds_bpermute`s from the hardware lane (`mbcnt`), like `get_sub_group_local_id`, rather than `activelane`. `ds_bpermute` only uses the low bits of the address, so a lane out of range reads an unspecified value instead of throwing, and the lane and offsets are truncated with `%` rather than converted with a check. They are only defined for the primitive types that `Device._shfl` decomposes into 32-bit shuffles (`Bool`, integers and IEEE floats), as is `supports_shuffle`, so that KernelInterface shuffles other `isbits` types, e.g. `Complex`, field by field. The votes are built on `ballot`, which only sets the bits of the active lanes, i.e. of the work-items of the sub-group. `get_max_sub_group_size` already is a constant of the generated code, as `fold_wavefrontsize!` folds `llvm.amdgcn.wavefrontsize` before optimization. Assisted-by: Claude Code (Opus 5.5)
Take KernelAbstractions and KernelInterface from the branch of JuliaGPU/KernelAbstractions.jl#831, which specifies the sub-group shuffles, votes and constant width, and the SPIRVIntrinsics it needs from OpenCL.jl's vc/subgroup-votes branch (JuliaGPU/OpenCL.jl#526), which KernelAbstractions gets from its [sources] that dependents don't pick up. Also allow AcceleratedKernels 0.5, which its main branch now is. Drop this commit once #831 is merged. Assisted-by: Claude Code (Opus 5.5)
KernelInterface now defines `shfl_down` and `shfl_up` to return the work-item's own value where the source lane is past the sub-group width (like CUDA's shuffles), rather than an unspecified value. `ds_bpermute` wraps the lane around, so select the work-item's own lane in that case, one compare and select with the constant wavefront size. The shuffles with a `width` use KernelInterface's fallbacks (a `ds_bpermute` from a computed lane), which is what `Device.shfl` etc. do too. Assisted-by: Claude Code (Opus 5.5)
The device overrides of `Base.min` and `Base.max` for floats called OCML's `__ocml_min`/`__ocml_max`, which, like C's `fmin`/`fmax`, return the other argument for a NaN, while Julia's `min` and `max` return NaN. Drop them: Julia's own definitions (`llvm.minimum`/`llvm.maximum` on Julia 1.12+, arithmetic before) compile for AMDGPU. Assisted-by: Claude Code (Opus 5.5)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the sub-group contract of KernelInterface from JuliaGPU/KernelAbstractions.jl#831 for
ROCBackend: the shufflesshfl,shfl_down,shfl_upandshfl_xor, the votessub_group_any,sub_group_allandsub_group_ballot, and aget_max_sub_group_sizethat is a constant of the generated code.What's implemented
shfl(val, lane)ds_bpermute(Device.bpermute) from hardware lanelane - 1shfl_down(val, offset)ds_bpermutefrom hardware lanembcnt + offset, ormbcntpast the wavefrontshfl_up(val, offset)ds_bpermutefrom hardware lanembcnt - offset, ormbcntif that is negativeshfl_xor(val, mask)ds_bpermutefrom hardware lanembcnt ⊻ masksub_group_ballot(pred)Device.ballot(pred)(llvm.amdgcn.ballot.i32/.i64, widened toUInt64)sub_group_any(pred)Device.ballot(pred) != 0sub_group_all(pred)Device.ballot(!pred) == 0get_max_sub_group_size()Device.wavefrontsize(), folded to a constant (unchanged)mbcnt), likeget_sub_group_local_id, notactivelane. They don't useDevice.shfl*, which clamp out-of-range offsets to the own lane and work onactivelane.ds_bpermuteonly uses the low bits of the address, so a lane out of range reads an unspecified value rather than trapping. Lanes and offsets are truncated with% Cint, so there is no checked conversion: the IR of a kernel using all shuffles and votes with runtimeIntlanes/offsets has nothrow/trap/unreachable.supports_shuffleare now only defined forconst ShuffleTypes = Union{Bool, Base.BitInteger, Base.IEEEFloat}(whatDevice._shfldecomposes into 32-bitbpermutes). The previous genericwhere {T}shfl_downoverride and thesupports_shufflethat listedComplexare gone, soComplexand otherisbitsstructs/tuples go through KernelInterface's field-by-field fallback.shfl_down/shfl_upreturn the work-item's own value where the source lane is past the wavefront (as KernelInterface now requires, like CUDA): one compare and select against the constant wavefront size before theds_bpermute.shfl_xorneeds no change, as KernelInterface requiresmaskbelow the width.width(KI.shfl(val, lane, width)etc.) use KernelInterface's fallbacks: oneds_bpermutefrom a lane computed with a few integer operations, the same asDevice.shfletc. DPP/ds_swizzlecould be cheaper for constant offsets, but isn't used here.ballotsetting only the bits of active lanes, i.e. of the work-items of the (possibly partial) sub-group. The result is uniform.Constant width
This needed no new code.
KI.kernel_functioncompiles for the device's wavefront size (and rejects a conflictingwavefrontsize64), andfold_wavefrontsize!infinish_module!already replacesllvm.amdgcn.wavefrontsizewith the compiled-for size before optimization. I checked this with@device_code_llvmon a kernel that storesKI.get_max_sub_group_size(): it compiles tostore i64 32withwavefrontsize64=falseandstore i64 64withwavefrontsize64=true, and the call tollvm.amdgcn.wavefrontsizeis gone. I added a comment saying so.[TEMP]commit[TEMP] Test against KernelAbstractions' vc/ki-subgroup-ops branchsits on top of the existing[TEMP]commit. It:vc/ki-subgroup-opsvc/subgroup-votes(lib/intrinsics, SPIRVIntrinsics: sub-group votes and collectives, unchecked shuffle lanes OpenCL.jl#526). KA gets it through its own[sources], which dependents don't pick up, so it is added to[extras]/[sources]here and developed on Julia 1.10 in the pipeline.mainnow is.Drop it once #831 is merged.
Note: KernelAbstractions
main, and so #831, require LLVM.jl 10 and GPUCompiler 2.10, butka-0.10is still on LLVM.jl 9. This branch (likeka-0.10against KAmain) won't resolve untilka-0.10picks up the LLVM.jl 10 port in #1132.Local testing
Update (after the shuffle/layout contract changes in JuliaGPU/KernelAbstractions.jl#831): KernelInterface testsuite: all pass but 4:
sub_group_reduce/sub_group_scanofmaxonFloat32with aNaN(32 and 29 work-items). That is not this PR: AMDGPU's device override ofBase.maxfor floats lowers tollvm.maxnum, which drops NaNs, unlike Julia'smax.kernelabstractions_tests: 2564 passed, 4 broken. wave32 and wave64 compile;shfl_down/shfl_upare av_cmp/v_cndmaskand ads_bpermuteeach.Earlier results:
Tested on an AMD Radeon RX 6800 XT (gfx1030, wave32), ROCm in
/opt/rocm, Julia 1.12.7. To get past the LLVM.jl 10 constraint, I used a local-only branch (not pushed): this branch merged withorigin/main, with the 5 commits oftb/llvm10(#1132) cherry-picked on top. The conflicts were resolved in favour ofka-0.10'sROCKernels.jl, local-memory initializer and wavefront fence.julia --project=test test/runtests.jl --jobs=2 kernelinterface kernelabstractions: 3152 passed, 4 broken, 0 failed (kernelinterface_testsandkernelabstractions_tests).Testsuite.testsuite(ROCBackend(), ROCArray)): 569/569 passed. This includes the new shuffle tests (lanes, out-of-range,shfl_up/shfl_xor, structs), votes and the width.@device_code_llvm/@device_code_gcnwithwavefrontsize64=falseandtrue: everything compiles, and the width is constant (see above). Only wave32 could be run on this GPU.Not tested locally: Julia 1.10/1.11, and running in wave64 mode.
🤖 Generated with Claude Code