Skip to content

KernelInterface - #3246

Closed
christiangnrd wants to merge 4 commits into
mainfrom
interface
Closed

christiangnrd wants to merge 4 commits into
mainfrom
interface

Conversation

@christiangnrd

@christiangnrd christiangnrd commented Aug 22, 2026 •

Copy link
Copy Markdown
Member

Do not merge until KernelInterface has been reviewed and interface fully decided

KernelInterface is the low-level, back-end-agnostic layer underneath KernelAbstractions 0.10: memory management, kernel compilation and launch with explicit workgroups, and device-side intrinsics (indices, barriers, local memory, sub-groups and shuffles, events). This PR implements it for CUDA, so code written against KernelInterface runs on NVIDIA GPUs.

The back-end is CUDACore.CUDAInterface, which KernelInterface.get_backend returns for a CuArray. It isn't public API yet. The existing KernelAbstractions back-end (CUDABackend) doesn't change here; #3302 ports it to KA 0.10 on top of this PR.

The implementation maps KernelInterface onto CUDA.jl's existing functionality: CuArray allocations, asynchronous copies on the task-local stream, kernel compilation and launch through cufunction, record_event/wait_event with CUDA events, and the indexing, sub-group and shuffle intrinsics with threadIdx/blockIdx/warp operations. KernelInterface.versioninfo forwards to CUDA.versioninfo, which now reports the KernelInterface version.

Tests: KernelInterface's testsuite, plus a CUDA test for partially filled warps.

@codecov

codecov Bot commented Aug 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 62.61682% with 40 lines in your changes missing coverage. Please review.
✅ Project coverage is 85.77%. Comparing base (f8606eb) to head (b9e867a).

Files with missing lines Patch % Lines
CUDACore/src/CUDAInterface.jl 61.90% 40 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #3246      +/-   ##
==========================================
- Coverage   85.89%   85.77%   -0.13%     
==========================================
  Files         187      188       +1     
  Lines       19022    19128     +106     
==========================================
+ Hits        16339    16407      +68     
- Misses       2683     2721      +38     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@christiangnrd
christiangnrd force-pushed the interface branch 3 times, most recently from ea8b69a to d670a43 Compare August 27, 2026 18:38
@christiangnrd
christiangnrd force-pushed the interface branch 3 times, most recently from 4e7f61c to 4dc5909 Compare September 9, 2026 00:09
@christiangnrd
christiangnrd force-pushed the interface branch 2 times, most recently from 2b942d2 to fdc35ad Compare September 21, 2026 21:44
@christiangnrd

Copy link
Copy Markdown
Member Author

I'm pretty sure the record_event and wait_event are implemented properly, but Fable says the test doesn't always fail when wait_event is disabled because an array on a different stream gets automatically synchronized

@christiangnrd
christiangnrd marked this pull request as ready for review September 22, 2026 20:17
@christiangnrd
christiangnrd force-pushed the interface branch 2 times, most recently from a5570fa to 561661d Compare September 23, 2026 12:36
Comment thread CUDACore/src/CUDACore.jl Outdated
Comment on lines +127 to +128
include("CUDAKernels.jl")
include("CUDAKernelsOld.jl")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is ugly?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. I found it was much easier in terms of rebasing to have the KI code in its future spot right away and then just turn the new file into the extension in the KA PR. I'm open to suggestions

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not have a single PR switching from the current interface to KernelInterface + KA.jl 0.10? Not that I particularly care, just curious.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I initially thought we could get KI merged independently before KA 0.10 so I had them split up but it might not turn out that way.

Comment thread CUDACore/src/CUDAKernels.jl Outdated
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: da561af Previous: c41ddb2 Ratio
array/accumulate/Float32/1d 100230 ns 99024 ns 1.01
array/accumulate/Float32/dims=1 71691 ns 72145 ns 0.99
array/accumulate/Float32/dims=1L 1587216 ns 1588461 ns 1.00
array/accumulate/Float32/dims=2 137617 ns 137848 ns 1.00
array/accumulate/Float32/dims=2L 655607 ns 654636 ns 1.00
array/accumulate/Int64/1d 118944 ns 117263 ns 1.01
array/accumulate/Int64/dims=1 76281 ns 76091 ns 1.00
array/accumulate/Int64/dims=1L 1698070 ns 1697362 ns 1.00
array/accumulate/Int64/dims=2 150344 ns 149255 ns 1.01
array/accumulate/Int64/dims=2L 987488 ns 986192 ns 1.00
array/broadcast 16632 ns 15428 ns 1.08
array/broadcast launch 8102 ns 6869 ns 1.18
array/construct 883.7234042553191 ns 870.7142857142857 ns 1.01
array/copy 16501 ns 16625 ns 0.99
array/copyto!/cpu_to_gpu 206117 ns 206218 ns 1.00
array/copyto!/gpu_to_cpu 240987 ns 239720 ns 1.01
array/copyto!/gpu_to_gpu 10250 ns 10239 ns 1.00
array/iteration/findall/bool 132970 ns 131631 ns 1.01
array/iteration/findall/int 144441 ns 143514 ns 1.01
array/iteration/findfirst/bool 69406 ns 68312 ns 1.02
array/iteration/findfirst/int 71491 ns 70287 ns 1.02
array/iteration/findmin/1d 64713 ns 60167 ns 1.08
array/iteration/findmin/2d 97830 ns 96150 ns 1.02
array/iteration/logical 187451 ns 178893 ns 1.05
array/iteration/scalar 63588 ns 63853 ns 1.00
array/permutedims/2d 46957 ns 45346 ns 1.04
array/permutedims/3d 48353 ns 46982 ns 1.03
array/permutedims/4d 49865 ns 47789 ns 1.04
array/random/rand/Float32 11509 ns 11572 ns 0.99
array/random/rand/Int64 19898 ns 18453 ns 1.08
array/random/rand!/Float32 7731.75 ns 7761.5 ns 1.00
array/random/rand!/Int64 16629 ns 16118 ns 1.03
array/random/randn/Float32 32460 ns 32223 ns 1.01
array/random/randn!/Float32 25186 ns 23621 ns 1.07
array/reductions/mapreduce/Float32/1d 32522 ns 32045 ns 1.01
array/reductions/mapreduce/Float32/dims=1 37401 ns 37283 ns 1.00
array/reductions/mapreduce/Float32/dims=1L 50997 ns 50954 ns 1.00
array/reductions/mapreduce/Float32/dims=2 55371 ns 54960 ns 1.01
array/reductions/mapreduce/Float32/dims=2L 67212 ns 66739 ns 1.01
array/reductions/mapreduce/Int64/1d 39695 ns 39254 ns 1.01
array/reductions/mapreduce/Int64/dims=1 40109 ns 39772 ns 1.01
array/reductions/mapreduce/Int64/dims=1L 88714 ns 88484 ns 1.00
array/reductions/mapreduce/Int64/dims=2 57709 ns 57117 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 83493 ns 82793 ns 1.01
array/reductions/reduce/Float32/1d 32722 ns 32298 ns 1.01
array/reductions/reduce/Float32/dims=1 37599 ns 36873 ns 1.02
array/reductions/reduce/Float32/dims=1L 50859 ns 50474 ns 1.01
array/reductions/reduce/Float32/dims=2 55329 ns 54877 ns 1.01
array/reductions/reduce/Float32/dims=2L 67374 ns 66929 ns 1.01
array/reductions/reduce/Int64/1d 40132 ns 38693 ns 1.04
array/reductions/reduce/Int64/dims=1 40145 ns 39824 ns 1.01
array/reductions/reduce/Int64/dims=1L 88619 ns 88360 ns 1.00
array/reductions/reduce/Int64/dims=2 57694 ns 57104 ns 1.01
array/reductions/reduce/Int64/dims=2L 83486 ns 82794 ns 1.01
array/reverse/1d 16889 ns 16811 ns 1.00
array/reverse/1dL 69642 ns 69626 ns 1.00
array/reverse/1dL_inplace 67326 ns 67260 ns 1.00
array/reverse/1d_inplace 8501.666666666666 ns 8336 ns 1.02
array/reverse/2d 19986 ns 19529 ns 1.02
array/reverse/2dL 73281 ns 73101 ns 1.00
array/reverse/2dL_inplace 67035 ns 67223 ns 1.00
array/reverse/2d_inplace 9681 ns 12107 ns 0.80
array/sorting/1d 2638703 ns 2646351 ns 1.00
array/sorting/2d 1013824 ns 1018059 ns 1.00
array/sorting/by 3158271 ns 3173715 ns 1.00
cuda/synchronization/context/auto 1016.6 ns 1015.9 ns 1.00
cuda/synchronization/context/blocking 780.8857142857142 ns 787.757281553398 ns 0.99
cuda/synchronization/context/nonblocking 5751.5 ns 5696.333333333333 ns 1.01
cuda/synchronization/stream/auto 897.7948717948718 ns 866.1666666666666 ns 1.04
cuda/synchronization/stream/blocking 682.4129032258064 ns 661.1446540880503 ns 1.03
cuda/synchronization/stream/nonblocking 5638.714285714285 ns 5718.666666666667 ns 0.99
integration/byval/reference 148091 ns 147659 ns 1.00
integration/byval/slices=1 149167 ns 148771 ns 1.00
integration/byval/slices=2 291741 ns 291271 ns 1.00
integration/byval/slices=3 435280 ns 434742 ns 1.00
integration/cudadevrt 105107 ns 104894 ns 1.00
integration/volumerhs 9145132 ns 9147760 ns 1.00
kernel/indexing 12902 ns 13026 ns 0.99
kernel/indexing_checked 13588 ns 13436 ns 1.01
kernel/launch 2138 ns 2109.3333333333335 ns 1.01
kernel/occupancy 710.6666666666666 ns 699.2582781456954 ns 1.02
kernel/rand 16367 ns 13938 ns 1.17
latency/import 4246104660 ns 4239937392 ns 1.00
latency/precompile 5017069479 ns 5025064161 ns 1.00
latency/ttfp 4714799305 ns 4726157535 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Comment thread test/core/kernelinterface.jl Outdated
Implement the KernelInterface API: allocation, asynchronous copies,
device selection, launch limits, kernel compilation and launch, the device-side
indexing, sub-group, local memory, barrier and shuffle intrinsics, and
`record_event`/`wait_event` using CUDA events.

The back-end lives in `CUDACore.CUDAInterface` and isn't public yet;
`KernelInterface.get_backend` returns it for a `CuArray`. The
KernelAbstractions back-end, `CUDABackend`, is unchanged.

`KernelInterface.versioninfo` forwards to `CUDA.versioninfo`, which now
also reports the KernelInterface version.

Co-authored-by: Tim Besard <tim.besard@gmail.com>
- implement `KI.launch` instead of the `Kernel` call method: KernelInterface now
  validates and sizes the launch
- split the launch limit (`max_work_group_size(kernel)`, the kernel's
  MAX_THREADS_PER_BLOCK) from the recommendation (`launch_configuration`, the
  occupancy API)
- compute the four primitive index queries with `% T`; KernelInterface derives
  the global ones
- typed sub-group queries; `get_sub_group_size` counts the threads of a partial
  last warp
- `supports_subgroups`/`supports_shuffle` instead of `shfl_down_types`, and
  report Float64 and atomics support, which no longer default to true
- `copyto!` checks lengths and only takes dense arrays, instead of copying
  `length(dst)` elements from `pointer(src)`
- `device(backend, A)`, `unsafe_free!`, `device!` returns nothing
- subtype `KI.Backend` and use the generic `zeros`/`ones`

KernelInterface 0.3 isn't registered yet: get it from the KernelAbstractions.jl
branch through `[sources]` (in every project depending on it, for
JuliaLang/Pkg.jl#4831), and develop it on Julia 1.10.
Julia doesn't specialize a method on `args...` that it only passes through, which
made every launch through the generic KernelInterface launch dispatch dynamically.
KernelInterface 0.3 passes the number of work-items to launch as `nitems`,
separately from the bound on the work-group size.
@maleadt

maleadt commented Sep 30, 2026

Copy link
Copy Markdown
Member

#3314

@maleadt maleadt closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants