KernelInterface - #474
Closed
christiangnrd wants to merge 6 commits into
Closed
KernelInterface#474christiangnrd wants to merge 6 commits into
christiangnrd wants to merge 6 commits into
Conversation
christiangnrd
force-pushed
the
kinterface
branch
from
August 22, 2026 15:32
dc740b9 to
0940095
Compare
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #474 +/- ##
==========================================
- Coverage 83.25% 83.04% -0.21%
==========================================
Files 19 20 +1
Lines 1618 1693 +75
==========================================
+ Hits 1347 1406 +59
- Misses 271 287 +16 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
christiangnrd
force-pushed
the
kinterface
branch
2 times, most recently
from
September 8, 2026 12:20
a1ef4c8 to
ad5cee5
Compare
christiangnrd
force-pushed
the
kinterface
branch
4 times, most recently
from
September 18, 2026 02:27
e602f87 to
4141e57
Compare
christiangnrd
marked this pull request as ready for review
September 22, 2026 19:29
christiangnrd
force-pushed
the
kinterface
branch
from
September 26, 2026 16:12
371952e to
b7c222a
Compare
Implement the KernelInterface (KI) API for OpenCL.jl in an internal `OpenCLInterface` module, reachable through `KI.get_backend` on a CLArray. The existing KernelAbstractions 0.9 back end moves unchanged to OpenCLKernelsOld.jl. Run the KI testsuite as part of the test suite. The device's work-group limits, including the per-dimension limits that KI 0.2.3 uses to bound automatically chosen workgroup sizes (NVIDIA, for example, only allows 64 work-items along the third dimension), are cached per device because querying them allocates on every launch. Co-authored-by: Tim Besard <tim.besard@gmail.com>
KernelInterface 0.3 lives on JuliaGPU/KernelAbstractions.jl branch tb/ki-0.3. Julia 1.11+ picks it up through `[sources]`, which has to be listed in every workspace project that depends on it (JuliaLang/Pkg.jl#4831). Julia 1.10 ignores `[sources]`, so it can't resolve until KernelInterface 0.3 is registered. Drop this commit then.
- subtype `KI.Backend`, and implement `KI.launch` instead of the kernel call method: KernelInterface now validates the launch and picks the sizes; - compute the typed index queries with `% T`, which has no error path, and let KernelInterface derive the global ones; the sub-group queries take a type too; - `max_work_group_size(kernel)` replaces `kernel_max_work_group_size`, and `max_num_groups` is only bounded by the global size; - report sub-groups only for devices where kernels get a fixed sub-group width (through `cl_intel_required_subgroup_size`), and shuffles for the types the device supports with `cl_khr_subgroup_shuffle` (replacing `shfl_down_types`); - conservative capabilities: float64 and unified memory per device, atomics; - `kernel_function` keeps the backend value it was given; - `copyto!` checks the lengths and returns the destination; forward `unsafe_free!`; - cache the device properties that launches and queries need per device. This needs SPIRVIntrinsics 1.1.3, whose 3-D builtins can be truncated without producing vector types that SPIR-V doesn't have.
The platform is part of an `OpenCLBackend`'s configuration, but it was only checked, with a warning, before allocating and launching, and `kernel_function` replaced it by the active platform. Instead, work for a backend whose platform isn't active now activates the platform's default device, as `KI.device!` would, so that the arrays it creates can be used afterwards, and host queries answer for that device. `get_backend` returns a backend for the array's platform. Kernel queries answer for the device the kernel was compiled for, rather than for the active one, and launching a kernel on another device throws an error instead of failing in the driver. Also implement `KI.device(backend, A)`.
`KI.synchronize` has to let other tasks run while it waits, but `clFinish` blocks the thread, and keeps the garbage collector from running on other threads. Like CUDA.jl, `cl.finish(queue; blocking=false)` now enqueues a marker and busy-waits for it briefly, then waits for it on a worker thread, in a GC-safe call, while the task yields. The other uses of `cl.finish` keep blocking: they include finalizers, which can't yield.
Julia doesn't specialize a method on `args...` that it only passes through, which made every launch through KernelInterface's generic launch dispatch dynamically (+1.2 µs and +1 kB per launch on CUDA).
Member
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.