Conversation
Add the sub-group operations that e.g. Molly.jl's CUDA kernels use, so that they can be written portably: - shuffles `shfl` (from a given lane), `shfl_up` and `shfl_xor`, next to `shfl_down`. Backends implement them for primitive types; a fallback shuffles `isbits` structs and tuples field by field, and `supports_shuffle` checks their fields. - votes `sub_group_any`, `sub_group_all` and `sub_group_ballot` (a `UInt64` mask, for sub-groups of at most 64 work-items), required with sub-group support. - `get_max_sub_group_size` is now required to be a constant of the generated code. Implement them for POCL; its sub-group width is folded into the IR before optimization. Assisted-by: Claude Code (Opus 5.5)
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
|
|
||
| # `sub_group_shuffle`, with the lane passed modulo `UInt32`, so that an out-of-range lane | ||
| # gives an unspecified value rather than an `InexactError` | ||
| for T in SPIRVIntrinsics.gentypes | ||
| @eval @device_function shuffle(x::$T, lane::Integer) = | ||
| @builtin_ccall( | ||
| "__spirv_GroupNonUniformShuffle", $T, (UInt32, $T, UInt32), | ||
| UInt32(Scope.Subgroup), x, (lane - 1) % UInt32 | ||
| ) | ||
| end | ||
|
|
||
| # Votes, from `cl_khr_subgroups` and `cl_khr_subgroup_ballot`. The SPIR-V back-end lowers | ||
| # these OpenCL built-ins, which have to be listed in `subgroup_intrinsics`. | ||
| const subgroup_intrinsics = ["_Z13sub_group_anyi", "_Z13sub_group_alli", "_Z16sub_group_balloti"] | ||
|
|
||
| @device_function sub_group_any(pred::Bool) = | ||
| ccall("extern _Z13sub_group_anyi", llvmcall, Int32, (Int32,), pred) != Int32(0) | ||
|
|
||
| @device_function sub_group_all(pred::Bool) = | ||
| ccall("extern _Z13sub_group_alli", llvmcall, Int32, (Int32,), pred) != Int32(0) | ||
|
|
||
| # bit `i` of the result is set for the lane with (0-based) id `i` | ||
| @device_function sub_group_ballot(pred::Bool) = | ||
| ccall("extern _Z16sub_group_balloti", llvmcall, NTuple{4, VecElement{UInt32}}, (Int32,), pred) |
There was a problem hiding this comment.
Should this be added to SPIRVIntrinsics.jl instad or are they hacks?
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #831 +/- ##
==========================================
- Coverage 79.06% 0.00% -79.07%
==========================================
Files 24 22 -2
Lines 2040 1857 -183
==========================================
- Hits 1613 0 -1613
- Misses 427 1857 +1430 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
On Julia 1.10, inference gives up on the recursive call of `shfl_fields` through the shuffle of a nested field (e.g. a tuple in a struct), leaving a dynamic invocation in the kernel. Generate the shuffles of all primitive fields directly instead. Assisted-by: Claude Code (Opus 5.5)
Replace the local workarounds with the votes and unchecked shuffle lanes from JuliaGPU/OpenCL.jl#526, taken from its branch until it is released: through `[sources]`, and explicitly where that doesn't apply (Julia 1.10 on CI, and the Buildkite jobs, whose OpenCL job developed SPIRVIntrinsics from OpenCL.jl's ka-0.10 branch). Assisted-by: Claude Code (Opus 5.5)
| # the sub-group width is fixed, so make `get_max_sub_group_size` a constant, as | ||
| # KernelInterface requires (this runs before optimization) | ||
| gvs = LLVM.globals(mod) | ||
| if haskey(gvs, "__spirv_BuiltInSubgroupMaxSize") | ||
| gv = gvs["__spirv_BuiltInSubgroupMaxSize"] | ||
| for use in collect(LLVM.uses(gv)) | ||
| load = LLVM.user(use) | ||
| load isa LLVM.LoadInst || continue | ||
| LLVM.replace_uses!(load, ConstantInt(LLVM.value_type(load), sg_size)) | ||
| LLVM.erase!(load) | ||
| end | ||
| isempty(LLVM.uses(gv)) && LLVM.erase!(gv) | ||
| end |
There was a problem hiding this comment.
This is perhaps a bit sketchy, and would need to be replicated for OpenCL/oneAPI
…ce and scan Fill the gaps that a survey of the packages using warp operations (KomaMRI, KernelIntrinsics/KernelForge, AcceleratedKernels#93, ParallelStencil, ClimaCore, ...) showed: - Primitive types that a backend doesn't support natively (e.g. `Bool`, `Char`, or 64-bit types on Metal) are shuffled as `UInt32` words, which backends now have to support. Structs keep being shuffled field by field. - Shuffles within segments of `width` lanes (`shfl(val, lane, width)` etc.), with CUDA's semantics, built on `shfl`. - `sub_group_match_any(val)`, the mask of the lanes with the same value, with a fallback built on `shfl` and `sub_group_ballot`. - `sub_group_reduce(op, val)` and `sub_group_scan(op, val)` with fallbacks built on the shuffles, which backends can implement with native operations. - Document how partial sub-groups behave. The new tests are in a function of their own: as part of `interface_testsuite`, compiling the host code crashed LLVM. Assisted-by: Claude Code (Opus 5.5)
KernelInterface (JuliaGPU/KernelAbstractions.jl#831) now shuffles primitive types that a back-end doesn't support natively as `UInt32` words, so the Metal-specific split into halves isn't needed anymore. Assisted-by: Claude Code (Opus 5.5)
Implement `KI.sub_group_reduce` and `KI.sub_group_scan` with the collectives of `cl_khr_subgroups` from SPIRVIntrinsics (JuliaGPU/OpenCL.jl#526): for `+` on 32- and 64-bit integers and floats, and `min`/`max` on integers. Floats keep the fallback for `min` and `max`, as OpenCL treats NaN and the sign of zero differently. Test the operators and types that backends may implement natively, including a NaN, and that POCL uses the native reduction. Assisted-by: Claude Code (Opus 5.5)
PoCL's `cl_khr_subgroups` reductions and scans lose the values of work-items that computed them in a divergent branch (PoCL 7.2; Intel's OpenCL runtime is fine), so `@groupreduce` in a `@kernel`, whose padding work-items are masked, returned garbage. Use KernelInterface's fallbacks again, and test reductions and scans of values from a divergent branch. Assisted-by: Claude Code (Opus 5.5)
Adds the sub-group primitives that Molly.jl's CUDA kernels use (
ext/MollyCUDAExt.jl), so that they can be written portably on top of KernelInterface. This is a companion to #830 (@groupreduce/@subgroupreduce).New device functions
shfl(val, lane)supports_shuffle(backend, T)shfl_up(val, offset)supports_shuffle(backend, T)shfl_xor(val, mask)supports_shuffle(backend, T)min/max/|) of tile maskssub_group_any(pred)/sub_group_all(pred)supports_subgroupsvote_any_sync)sub_group_ballot(pred)::UInt64supports_subgroups, width ≤ 64count_ones/trailing_zeroson the resultisbitsstructs and tuples field by field (e.g.SVector, Unitful quantities). For such types,supports_shufflechecks their fields. Molly currently gets this by overriding CUDA.jl's internalshfl_recurse.supports_shuffleand the backend implementation notes now cover all four shuffles.Warp size
get_max_sub_group_size()is now required to be a compile-time constant of the generated code. Loops over lanes and shuffle butterflies get specialized for it, which every backend can provide since kernels are already compiled for a fixed width (sub_group_size(backend)). For type-level decisions (e.g. aUInt32vsUInt64mask) the docs point to passingsub_group_size(backend)from the host.POCL
sub_group_shuffle/sub_group_shuffle_xor. Out-of-range lanes give an unspecified value instead of anInexactError.sub_group_any/sub_group_all/sub_group_ballot(cl_khr_subgroups,cl_khr_subgroup_ballot).finish_module!replaces loads of__spirv_BuiltInSubgroupMaxSizewith the width the kernel is compiled for (intel_reqd_sub_group_size), before optimization.Depends on JuliaGPU/OpenCL.jl#526, which adds the votes and unchecked shuffle lanes to SPIRVIntrinsics. Until it's released, this PR takes SPIRVIntrinsics from that branch:
[sources]inProject.toml;[sources];ka-0.10branch.The branch has the LLVM 10 upgrade that
ka-0.10lacks. Before merging, these should be replaced by a compat bound on the release; all of them are markedTODO.Built on the above (fallbacks, backends may override)
Added after a survey of packages that use warp operations: KomaMRI, KernelIntrinsics/KernelForge, AcceleratedKernels#93, ParallelStencil, ClimaCore, IntervalMDP, …
Bool,Char, 64-bit types on Metal, …)UInt32words. Backends now have to supportUInt32natively.shfl(val, lane, width),shfl_down/up/xor(val, x, width)widthlanes with CUDA's semantics (reads from outside the segment return the own value), built onshflsub_group_match_any(val)::UInt64===), found group by group withshfl+sub_group_ballot. CUDA could usematch.any.syncsub_group_reduce(op, val)shfl_down, then broadcast withshfl. Backends can dispatch ontypeof(op)for native reductions (Metalsimd_sum, SPIR-VGroupNonUniformIAdd, CUDAredux.sync)sub_group_scan(op, val)shfl_upThe docs now also say how partial sub-groups behave: lanes without a work-item give unspecified shuffle values, and the votes, match, reduce and scan only take the existing work-items into account.
The new tests are in small functions of their own,
subgroup_communication_testsuiteand helpers, with concrete loops. An earlier version iterated over heterogeneous tuples of functions and passed the resulting union-of-singletons value throughKI.@launch, whoseGC.@preservethen hit JuliaLang/julia#63482: on 1.12 and 1.13, codegen emits a nullgc_preserve_beginoperand, and LLVM segfaults. That is fixed on master by #63483, but the backport to 1.12/1.13 is still pending.For backend packages
KernelInterface 0.4 is still unreleased, so this extends its contract. CUDA.jl, AMDGPU.jl, oneAPI.jl, Metal.jl and OpenCL.jl will need to:
shfl,shfl_up,shfl_xor,sub_group_any,sub_group_allandsub_group_ballot;shfl_downoverrides to the primitive types, so that struct shuffles reach the fallback;get_max_sub_group_size(e.g.32 % Ton CUDA).One open question:
sub_group_ballotis required for widths ≤ 64. If a backend can't support ballot, it could get its own capability query instead.Tests
shfl,shfl_uplanes and ashfl_xorbutterfly all-reduce for every supported typesupports_shufflefor structsany/all/ballotwith several predicate patternsget_max_sub_group_sizereturns the width🤖 Generated with Claude Code