KernelInterface - #3246
KernelInterface#3246christiangnrd wants to merge 4 commits into
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #3246 +/- ##
==========================================
- Coverage 85.89% 85.77% -0.13%
==========================================
Files 187 188 +1
Lines 19022 19128 +106
==========================================
+ Hits 16339 16407 +68
- Misses 2683 2721 +38 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
ea8b69a to
d670a43
Compare
4e7f61c to
4dc5909
Compare
4dc5909 to
dbe0b1e
Compare
2b942d2 to
fdc35ad
Compare
|
I'm pretty sure the record_event and wait_event are implemented properly, but Fable says the test doesn't always fail when wait_event is disabled because an array on a different stream gets automatically synchronized |
ce57720 to
fdc35ad
Compare
a5570fa to
561661d
Compare
| include("CUDAKernels.jl") | ||
| include("CUDAKernelsOld.jl") |
There was a problem hiding this comment.
Yes. I found it was much easier in terms of rebasing to have the KI code in its future spot right away and then just turn the new file into the extension in the KA PR. I'm open to suggestions
There was a problem hiding this comment.
Why not have a single PR switching from the current interface to KernelInterface + KA.jl 0.10? Not that I particularly care, just curious.
There was a problem hiding this comment.
I initially thought we could get KI merged independently before KA 0.10 so I had them split up but it might not turn out that way.
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
446c439 to
744314f
Compare
Implement the KernelInterface API: allocation, asynchronous copies, device selection, launch limits, kernel compilation and launch, the device-side indexing, sub-group, local memory, barrier and shuffle intrinsics, and `record_event`/`wait_event` using CUDA events. The back-end lives in `CUDACore.CUDAInterface` and isn't public yet; `KernelInterface.get_backend` returns it for a `CuArray`. The KernelAbstractions back-end, `CUDABackend`, is unchanged. `KernelInterface.versioninfo` forwards to `CUDA.versioninfo`, which now also reports the KernelInterface version. Co-authored-by: Tim Besard <tim.besard@gmail.com>
- implement `KI.launch` instead of the `Kernel` call method: KernelInterface now validates and sizes the launch - split the launch limit (`max_work_group_size(kernel)`, the kernel's MAX_THREADS_PER_BLOCK) from the recommendation (`launch_configuration`, the occupancy API) - compute the four primitive index queries with `% T`; KernelInterface derives the global ones - typed sub-group queries; `get_sub_group_size` counts the threads of a partial last warp - `supports_subgroups`/`supports_shuffle` instead of `shfl_down_types`, and report Float64 and atomics support, which no longer default to true - `copyto!` checks lengths and only takes dense arrays, instead of copying `length(dst)` elements from `pointer(src)` - `device(backend, A)`, `unsafe_free!`, `device!` returns nothing - subtype `KI.Backend` and use the generic `zeros`/`ones` KernelInterface 0.3 isn't registered yet: get it from the KernelAbstractions.jl branch through `[sources]` (in every project depending on it, for JuliaLang/Pkg.jl#4831), and develop it on Julia 1.10.
Julia doesn't specialize a method on `args...` that it only passes through, which made every launch through the generic KernelInterface launch dispatch dynamically.
KernelInterface 0.3 passes the number of work-items to launch as `nitems`, separately from the bound on the work-group size.
Do not merge until KernelInterface has been reviewed and interface fully decided
KernelInterface is the low-level, back-end-agnostic layer underneath KernelAbstractions 0.10: memory management, kernel compilation and launch with explicit workgroups, and device-side intrinsics (indices, barriers, local memory, sub-groups and shuffles, events). This PR implements it for CUDA, so code written against KernelInterface runs on NVIDIA GPUs.
The back-end is
CUDACore.CUDAInterface, whichKernelInterface.get_backendreturns for aCuArray. It isn't public API yet. The existing KernelAbstractions back-end (CUDABackend) doesn't change here; #3302 ports it to KA 0.10 on top of this PR.The implementation maps KernelInterface onto CUDA.jl's existing functionality:
CuArrayallocations, asynchronous copies on the task-local stream, kernel compilation and launch throughcufunction,record_event/wait_eventwith CUDA events, and the indexing, sub-group and shuffle intrinsics withthreadIdx/blockIdx/warp operations.KernelInterface.versioninfoforwards toCUDA.versioninfo, which now reports the KernelInterface version.Tests: KernelInterface's testsuite, plus a CUDA test for partially filled warps.