Repository navigation
Conversation
KernelAbstractions 0.10 builds on KernelInterface, which defines what a back end provides: memory and device management, compiling and launching kernels, and the device-side intrinsics. KernelAbstractions then launches `@kernel` kernels itself on any KernelInterface back end, which replaces Metal's copy of that launch path: partitioning the ndrange, building the kernel's context, and sizing the threadgroups. `MetalBackend` now implements KernelInterface, and Metal depends on it instead of on KernelAbstractions. What KernelAbstractions still needs from a back end moves to an extension: the stack allocation behind `@private`, and the Adapt rule for moving arrays to the CPU. KernelInterface's default `launch_configuration` uses the pipeline's `maxTotalThreadsPerThreadgroup`, as Metal's launch did. `KI.kernel_function` receives the callable unconverted, and keeps it as the kernel's source, which is converted again at every launch, as with `@metal`: the buffers a closure captures are declared to the encoder and kept alive. `KI.launch` passes the kernel arguments on as a tuple, to the same launch function `HostKernel` calls, so kernels with many arguments aren't splatted. It passes Metal's `queue` and `submit` options on, and rejects `threads` and `groups`, which would override the launch geometry that KernelInterface validated. `KI.copyto!` also accepts contiguous views of host arrays, and `adapt(MetalBackend(), x)` moves any array to the GPU, as `adapt(MtlArray, x)` does, rather than only `Array`s. KernelAbstractions now computes indices in 32 bits where they fit, so its conversion tests pass and are no longer skipped. Its test of a kernel with 41 arguments is skipped instead, since a Metal kernel takes at most 31 buffers. Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Sub-groups are SIMD-groups: implement the sub-group queries, `sub_group_barrier` and `shfl_down`, and report their support to KernelInterface, with `shfl_down` for the types `simd_shuffle_down` supports. The sub-group width that `KI.sub_group_size` promises is 32, which `kernel_function` checks against the pipeline's `threadExecutionWidth`. KernelInterface leaves unspecified how work-items are grouped into sub-groups; Metal forms SIMD-groups from consecutive linear thread indices, so the last one of a threadgroup can be partial, which a test checks. Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Implement `KI.record_event` with an `MTLSharedEvent` that a command buffer signals after the task's queued work, and `KI.wait_event` by encoding a wait for it in the task's open batch of work, so that the GPU waits instead of the host. `KernelAbstractions.@spawn` uses them to order a new task's work after the work its parent had queued, without synchronizing the parent; the default is a full synchronization. KernelInterface's testsuite checks them once `record_event` returns an event. Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
`KI.versioninfo(MetalBackend())` prints `Metal.versioninfo()`.
… main branch Neither is registered yet. Julia 1.10 ignores [sources] and Julia 1.11 doesn't support workspaces, so Buildkite sets up the test environment by hand there, and runs the tests in it directly. Drop this commit once both are registered.
Implement the sub-group communication contract of KernelInterface (JuliaGPU/KernelAbstractions.jl#831): - `shfl`, `shfl_down`, `shfl_up` and `shfl_xor` map to the SIMD-group shuffles, only for the primitive types Metal supports, so that other `isbits` types reach KernelInterface's fallback that shuffles them field by field. `supports_shuffle` is restricted likewise. Lanes, offsets and masks are wrapped with `% Int16` instead of converted, so that out of range they give an unspecified value rather than throwing, and lanes and masks are reduced to the SIMD-group width. - `Int64` and `UInt64` are shuffled as two 32-bit halves, since Metal has no 64-bit shuffles. - `sub_group_ballot` is `simd_ballot`, and `sub_group_any`/`sub_group_all` are derived from it (inactive lanes' bits are clear, so `all` checks the ballot of `!pred`). - `get_max_sub_group_size` returns the constant `SIMD_WIDTH`, which `kernel_function` already checks the pipeline's execution width against, so that code depending on it is specialized for it. Assisted-by: Claude Code (Opus 5.5)
Get KernelAbstractions and KernelInterface from KernelAbstractions.jl's `vc/ki-subgroup-ops` branch instead of `main`, and SPIRVIntrinsics, which that branch needs, from OpenCL.jl's `vc/subgroup-votes` branch (JuliaGPU/OpenCL.jl#526). Julia 1.10 ignores [sources], so Buildkite checks out and develops all three there. Drop this commit once JuliaGPU/KernelAbstractions.jl#831 is merged. Assisted-by: Claude Code (Opus 5.5)
KernelInterface (JuliaGPU/KernelAbstractions.jl#831) now shuffles primitive types that a back-end doesn't support natively as `UInt32` words, so the Metal-specific split into halves isn't needed anymore. Assisted-by: Claude Code (Opus 5.5)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the sub-group communication contract that JuliaGPU/KernelAbstractions.jl#831 adds to KernelInterface, on top of
ka-0.10.What's implemented
shfl(val, lane)simd_shuffle(val, lane)shfl_down(val, offset)simd_shuffle_down(val, offset)shfl_up(val, offset)simd_shuffle_up(val, offset)shfl_xor(val, mask)simd_shuffle_xor(val, mask)sub_group_ballot(pred)simd_ballot(pred)sub_group_any(pred)simd_ballot(pred) != 0sub_group_all(pred)simd_ballot(!pred) == 0get_max_sub_group_size(T)SIMD_WIDTH % T(wasthreads_per_simdgroup() % T)supports_shuffleonly cover the primitive types Metal supports (Float32, Float16, (U)Int32, (U)Int16, (U)Int8, plus 64-bit integers, see below). Otherisbitstypes reach KernelInterface's fallback, which shuffles structs and tuples field by field. Before this PR,shfl_downwas defined for everyTandsupports_shufflereturnedfalsefor composites.Int16, and accallconversion is checked. The overrides wrap lanes, offsets and masks with% Int16. Lanes and masks are also reduced modulo the SIMD-group width, because MSL requiressimd_shuffleto get a valid lane id. Out-of-range values give an unspecified value, as the contract allows.Int64andUInt64are split into twoUInt32halves, each half goes through the 32-bit intrinsic, and the halves are recombined. A comment notes that this could move into KernelInterface's fallback later.Float64is not included: Metal kernels can't use it at all (supports_float64isfalse).sub_group_all.simd_ballotclears the bits of inactive lanes, such as those past the end of a partial SIMD-group. Checking that the ballot of!predis zero therefore only looks at active lanes, without relying on the exact semantics ofsimd_vote_all.get_max_sub_group_sizeto be a compile-time constant of the generated code.KI.kernel_functionalready errors unless the pipeline'sthreadExecutionWidth == SIMD_WIDTH(32), so returning the constantSIMD_WIDTHis safe for every kernel launched through KernelInterface.[TEMP]commitThe second commit points the
[sources]for KernelAbstractions and KernelInterface at KernelAbstractions.jl'svc/ki-subgroup-opsbranch instead ofmain. It also adds SPIRVIntrinsics from OpenCL.jl'svc/subgroup-votesbranch (JuliaGPU/OpenCL.jl#526), which that KA branch needs, to the test project. On Julia 1.10, Buildkite checks out all three and develops them, so CI tests against #831. Drop the commit once #831 is merged.Verification
Untested on Apple hardware. This was written on Linux, where Metal.jl's runtime can't load. The Buildkite Metal jobs, which run KernelInterface's sub-group testsuite (
test/kernelinterface.jl), will be the first real test.Done on the host:
src/MetalKernels.jl:Int64andUInt64, including extremes and random values.Int16conversion throw. In-range lanes and masks are unchanged.supports_shuffleis true for the native types and(U)Int64.Float64,BoolandRef.ShuffleStruct-style struct with anInt64field.Int64field.Float64reaches the fallback'sArgumentError.Not done: GPU compilation (AIR/metallib validation) and execution.
🤖 Generated with Claude Code