Skip to content

KernelInterface: sub-group shuffles, votes and a constant width - #998

Draft
vchuravy wants to merge 8 commits into
ka-0.10from
vc/ki-subgroup-ops
Draft

vchuravy wants to merge 8 commits into
ka-0.10from
vc/ki-subgroup-ops

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026

Copy link
Copy Markdown
Member

Implements the sub-group communication contract that JuliaGPU/KernelAbstractions.jl#831 adds to KernelInterface, on top of ka-0.10.

What's implemented

KernelInterface Metal
shfl(val, lane) simd_shuffle(val, lane)
shfl_down(val, offset) simd_shuffle_down(val, offset)
shfl_up(val, offset) simd_shuffle_up(val, offset)
shfl_xor(val, mask) simd_shuffle_xor(val, mask)
sub_group_ballot(pred) simd_ballot(pred)
sub_group_any(pred) simd_ballot(pred) != 0
sub_group_all(pred) simd_ballot(!pred) == 0
get_max_sub_group_size(T) SIMD_WIDTH % T (was threads_per_simdgroup() % T)
  • Restricted signatures. The shuffle overrides and supports_shuffle only cover the primitive types Metal supports (Float32, Float16, (U)Int32, (U)Int16, (U)Int8, plus 64-bit integers, see below). Other isbits types reach KernelInterface's fallback, which shuffles structs and tuples field by field. Before this PR, shfl_down was defined for every T and supports_shuffle returned false for composites.
  • Out-of-range arguments don't throw. The intrinsics take an Int16, and a ccall conversion is checked. The overrides wrap lanes, offsets and masks with % Int16. Lanes and masks are also reduced modulo the SIMD-group width, because MSL requires simd_shuffle to get a valid lane id. Out-of-range values give an unspecified value, as the contract allows.
  • 64-bit shuffles. Metal has no 64-bit shuffles. Int64 and UInt64 are split into two UInt32 halves, each half goes through the 32-bit intrinsic, and the halves are recombined. A comment notes that this could move into KernelInterface's fallback later. Float64 is not included: Metal kernels can't use it at all (supports_float64 is false).
  • sub_group_all. simd_ballot clears the bits of inactive lanes, such as those past the end of a partial SIMD-group. Checking that the ballot of !pred is zero therefore only looks at active lanes, without relying on the exact semantics of simd_vote_all.
  • Constant width. KernelInterface now requires get_max_sub_group_size to be a compile-time constant of the generated code. KI.kernel_function already errors unless the pipeline's threadExecutionWidth == SIMD_WIDTH (32), so returning the constant SIMD_WIDTH is safe for every kernel launched through KernelInterface.

[TEMP] commit

The second commit points the [sources] for KernelAbstractions and KernelInterface at KernelAbstractions.jl's vc/ki-subgroup-ops branch instead of main. It also adds SPIRVIntrinsics from OpenCL.jl's vc/subgroup-votes branch (JuliaGPU/OpenCL.jl#526), which that KA branch needs, to the test project. On Julia 1.10, Buildkite checks out all three and develops them, so CI tests against #831. Drop the commit once #831 is merged.

Verification

Untested on Apple hardware. This was written on Linux, where Metal.jl's runtime can't load. The Buildkite Metal jobs, which run KernelInterface's sub-group testsuite (test/kernelinterface.jl), will be the first real test.

Done on the host:

  • Every changed file parses, and the TOML and YAML files load.
  • The helpers, extracted verbatim from src/MetalKernels.jl:
    • The 64-bit split/join round-trips for Int64 and UInt64, including extremes and random values.
    • An emulated 32-lane shuffle (rotate, xor, reverse) returns the source lane's 64-bit value.
    • Lane and mask wrapping never makes the emulated intrinsic's checked Int16 conversion throw. In-range lanes and masks are unchanged.
    • The vote formulas are correct for emulated partial SIMD-groups.
  • Dispatch against KernelInterface from Remove GPU specification from Benchmark run #831, with the overrides evaluated as plain methods and the intrinsics stubbed:
    • supports_shuffle is true for the native types and (U)Int64.
    • It is false for Float64, Bool and Ref.
    • It is true for KI's ShuffleStruct-style struct with an Int64 field.
    • A struct shuffle hits the 32-bit intrinsic once per primitive field and twice for the Int64 field.
    • Float64 reaches the fallback's ArgumentError.

Not done: GPU compilation (AIR/metallib validation) and execution.

🤖 Generated with Claude Code

maleadt and others added 8 commits September 30, 2026 19:29
KernelAbstractions 0.10 builds on KernelInterface, which defines what a back
end provides: memory and device management, compiling and launching kernels,
and the device-side intrinsics. KernelAbstractions then launches `@kernel`
kernels itself on any KernelInterface back end, which replaces Metal's copy of
that launch path: partitioning the ndrange, building the kernel's context, and
sizing the threadgroups.

`MetalBackend` now implements KernelInterface, and Metal depends on it instead
of on KernelAbstractions. What KernelAbstractions still needs from a back end
moves to an extension: the stack allocation behind `@private`, and the Adapt
rule for moving arrays to the CPU. KernelInterface's default
`launch_configuration` uses the pipeline's `maxTotalThreadsPerThreadgroup`, as
Metal's launch did.

`KI.kernel_function` receives the callable unconverted, and keeps it as the
kernel's source, which is converted again at every launch, as with `@metal`:
the buffers a closure captures are declared to the encoder and kept alive.

`KI.launch` passes the kernel arguments on as a tuple, to the same launch
function `HostKernel` calls, so kernels with many arguments aren't splatted.
It passes Metal's `queue` and `submit` options on, and rejects `threads` and
`groups`, which would override the launch geometry that KernelInterface
validated. `KI.copyto!` also accepts contiguous views of host arrays, and
`adapt(MetalBackend(), x)` moves any array to the GPU, as `adapt(MtlArray, x)`
does, rather than only `Array`s.

KernelAbstractions now computes indices in 32 bits where they fit, so its
conversion tests pass and are no longer skipped. Its test of a kernel with 41
arguments is skipped instead, since a Metal kernel takes at most 31 buffers.

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Sub-groups are SIMD-groups: implement the sub-group queries,
`sub_group_barrier` and `shfl_down`, and report their support to
KernelInterface, with `shfl_down` for the types `simd_shuffle_down` supports.
The sub-group width that `KI.sub_group_size` promises is 32, which
`kernel_function` checks against the pipeline's `threadExecutionWidth`.

KernelInterface leaves unspecified how work-items are grouped into sub-groups;
Metal forms SIMD-groups from consecutive linear thread indices, so the last one
of a threadgroup can be partial, which a test checks.

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Implement `KI.record_event` with an `MTLSharedEvent` that a command buffer
signals after the task's queued work, and `KI.wait_event` by encoding a wait for
it in the task's open batch of work, so that the GPU waits instead of the host.
`KernelAbstractions.@spawn` uses them to order a new task's work after the work
its parent had queued, without synchronizing the parent; the default is a full
synchronization.

KernelInterface's testsuite checks them once `record_event` returns an event.

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
`KI.versioninfo(MetalBackend())` prints `Metal.versioninfo()`.
… main branch

Neither is registered yet. Julia 1.10 ignores [sources] and Julia 1.11 doesn't
support workspaces, so Buildkite sets up the test environment by hand there,
and runs the tests in it directly.

Drop this commit once both are registered.
Implement the sub-group communication contract of KernelInterface
(JuliaGPU/KernelAbstractions.jl#831):

- `shfl`, `shfl_down`, `shfl_up` and `shfl_xor` map to the SIMD-group shuffles, only for
  the primitive types Metal supports, so that other `isbits` types reach KernelInterface's
  fallback that shuffles them field by field. `supports_shuffle` is restricted likewise.
  Lanes, offsets and masks are wrapped with `% Int16` instead of converted, so that out of
  range they give an unspecified value rather than throwing, and lanes and masks are
  reduced to the SIMD-group width.
- `Int64` and `UInt64` are shuffled as two 32-bit halves, since Metal has no 64-bit
  shuffles.
- `sub_group_ballot` is `simd_ballot`, and `sub_group_any`/`sub_group_all` are derived from
  it (inactive lanes' bits are clear, so `all` checks the ballot of `!pred`).
- `get_max_sub_group_size` returns the constant `SIMD_WIDTH`, which `kernel_function`
  already checks the pipeline's execution width against, so that code depending on it is
  specialized for it.

Assisted-by: Claude Code (Opus 5.5)
Get KernelAbstractions and KernelInterface from KernelAbstractions.jl's
`vc/ki-subgroup-ops` branch instead of `main`, and SPIRVIntrinsics, which that branch
needs, from OpenCL.jl's `vc/subgroup-votes` branch (JuliaGPU/OpenCL.jl#526). Julia 1.10
ignores [sources], so Buildkite checks out and develops all three there.

Drop this commit once JuliaGPU/KernelAbstractions.jl#831 is merged.

Assisted-by: Claude Code (Opus 5.5)
KernelInterface (JuliaGPU/KernelAbstractions.jl#831) now shuffles primitive
types that a back-end doesn't support natively as `UInt32` words, so the
Metal-specific split into halves isn't needed anymore.

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants