Conversation
This was referenced Sep 30, 2026
Closed
KernelAbstractions 0.10 builds on KernelInterface, which defines what a back end provides: memory and device management, compiling and launching kernels, and the device-side intrinsics. KernelAbstractions then launches `@kernel` kernels itself on any KernelInterface back end, which replaces CUDA's copy of that launch path: partitioning the ndrange, building the kernel's context, and tuning the workgroup size. `CUDABackend` now implements KernelInterface in CUDACore, which depends on it instead of on KernelAbstractions. What KernelAbstractions still needs from a back end moves to an extension: the `MArray` behind `@private`, the Adapt rule for `@Const`, and the `maxthreads` hint for a static workgroup size. `prefer_blocks` now applies in `KI.launch_configuration`, which receives the number of work-items. `KI.launch` passes the kernel arguments on as a tuple, so kernels with many arguments stay cheap to launch, as #3309 made them for the old launch path. It rejects `threads` and `blocks`, which would override the launch geometry that KernelInterface validated. `KI.copyto!` accepts dense arrays and contiguous views of them, as the old `KA.copyto!` did. KernelAbstractions converts the arguments twice, to determine the types to compile for and again when launching, where CUDA's launch converted them once. `cudaconvert` has to be pure, so this only shows with conversions that have side effects, as in two tests that count them. Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Sub-groups are warps: implement the sub-group queries, `sub_group_barrier` and `shfl_down`, and report their support to KernelInterface. KernelInterface leaves unspecified how work-items are grouped into sub-groups; CUDA forms warps from consecutive linear thread indices, so the last warp of a block can be partial, which a test checks. Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
Implement `KI.record_event` and `KI.wait_event` with a `CuEvent` recorded on, and waited for by, the task's stream. `KernelAbstractions.@spawn` uses them to order a new task's work after the work its parent had queued, without synchronizing the parent; the default is a full synchronization. KernelInterface's testsuite checks them once `record_event` returns an event.
`KI.versioninfo(CUDABackend())` prints `CUDA.versioninfo()`.
… main branch Neither is registered yet. Julia 1.10 and 1.11 don't pick up the test project's [sources], so Buildkite develops both explicitly there, and the GPU-less and Enzyme jobs move to Julia 1.12. Drop this commit once both are registered.
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
This was referenced Oct 2, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
KernelAbstractions 0.10 moves its back-end API into a separate package, KernelInterface. A back end implements KernelInterface: allocating and copying memory, selecting devices, compiling and launching kernels, and the device-side intrinsics. KernelAbstractions then launches
@kernelkernels the same way on every back end (JuliaGPU/KernelAbstractions.jl#801). This PR ports CUDA.jl to that model in one step. It replaces #3246, which added KernelInterface next to the existing KernelAbstractions 0.9 back end, and #3302/#3304, which then ported that to KernelAbstractions 0.10.CUDABackendnow implements KernelInterface, in CUDACore, which depends on KernelInterface instead of KernelAbstractions. That makes the KernelInterface layer usable on its own, without KernelAbstractions' macros:KernelAbstractions becomes a weak dependency. Its extension only provides what KernelAbstractions still needs from a back end: the
MArraybehind@private, the Adapt rule behind@Const, and themaxthreadshint for kernels with a static workgroup size. CUDA's own copy of the KernelAbstractions launch path goes away: partitioning the ndrange, building the kernel's context, tuning the workgroup size, and calling the kernel. What remains specific to KernelAbstractions is a 37-line extension.Kernels behave the same, but they are launched differently. KernelAbstractions now launches them on a 3-D grid, and computes
@indexin 32 bits when the iteration space fits (JuliaGPU/KernelAbstractions.jl#797). So a kernel over a 3-Dndrangedoesn't need divisions to compute its index. The test suite checks that for the generated PTX. With a KA kernel that copies aFloat32array, on an RTX 5080:@index(Global, Cartesian)@index(Global, Linear)@index(Global, Cartesian)@index(Global, Linear)CUDABackend(; prefer_blocks=true)still prefers more, smaller blocks. It now takes effect inKI.launch_configuration, which receives the number of work-items to cover.Launching kernels with many arguments stays as cheap as with #3309. KernelInterface 0.4 passes the arguments to the back end as a tuple (JuliaGPU/KernelAbstractions.jl#811), and CUDA forwards that tuple to its own launch.
KI.launchrejectsthreadsandblocks, which would override the launch geometry KernelInterface has validated, but passes CUDA's other launch options on:The first commit is the port. The next ones implement parts of KernelInterface that are optional: sub-groups (warps, including a partial last warp), ordering work across tasks with CUDA events for
KernelAbstractions.@spawn, andKI.versioninfo.One regression remains. KernelAbstractions converts the arguments twice, once to determine the argument types to compile for and once when launching, where CUDA converted them once.
cudaconvertis supposed to be pure, so this only shows with a conversion that has side effects. Two tests that count conversions are now@test_broken.This needs a breaking release, since it requires KernelAbstractions 0.10, which isn't registered yet. Until then, the last commit takes KernelAbstractions and KernelInterface from their development branch, through
[sources]and, for Julia 1.10 and 1.11, which don't pick those up, by developing them explicitly on Buildkite. That commit is dropped before merging.