Skip to content

KernelInterface - #474

Closed
christiangnrd wants to merge 6 commits into
mainfrom
kinterface
Closed

christiangnrd wants to merge 6 commits into
mainfrom
kinterface

Conversation

@christiangnrd

Copy link
Copy Markdown
Member

No description provided.

@codecov

codecov Bot commented Aug 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.08280% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 83.04%. Comparing base (eb27780) to head (24aa1f2).

Files with missing lines Patch % Lines
src/OpenCLKernelsOld.jl 90.47% 10 Missing ⚠️
src/OpenCLKernels.jl 92.30% 4 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #474      +/-   ##
==========================================
- Coverage   83.25%   83.04%   -0.21%     
==========================================
  Files          19       20       +1     
  Lines        1618     1693      +75     
==========================================
+ Hits         1347     1406      +59     
- Misses        271      287      +16     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@christiangnrd
christiangnrd force-pushed the kinterface branch 2 times, most recently from a1ef4c8 to ad5cee5 Compare September 8, 2026 12:20
@christiangnrd
christiangnrd force-pushed the kinterface branch 4 times, most recently from e602f87 to 4141e57 Compare September 18, 2026 02:27
@christiangnrd
christiangnrd marked this pull request as ready for review September 22, 2026 19:29
Implement the KernelInterface (KI) API for OpenCL.jl in an internal
`OpenCLInterface` module, reachable through `KI.get_backend` on a CLArray.
The existing KernelAbstractions 0.9 back end moves unchanged to
OpenCLKernelsOld.jl. Run the KI testsuite as part of the test suite.

The device's work-group limits, including the per-dimension limits that
KI 0.2.3 uses to bound automatically chosen workgroup sizes (NVIDIA, for
example, only allows 64 work-items along the third dimension), are
cached per device because querying them allocates on every launch.

Co-authored-by: Tim Besard <tim.besard@gmail.com>
KernelInterface 0.3 lives on JuliaGPU/KernelAbstractions.jl branch tb/ki-0.3.
Julia 1.11+ picks it up through `[sources]`, which has to be listed in every
workspace project that depends on it (JuliaLang/Pkg.jl#4831). Julia 1.10
ignores `[sources]`, so it can't resolve until KernelInterface 0.3 is
registered. Drop this commit then.
- subtype `KI.Backend`, and implement `KI.launch` instead of the kernel call
  method: KernelInterface now validates the launch and picks the sizes;
- compute the typed index queries with `% T`, which has no error path, and let
  KernelInterface derive the global ones; the sub-group queries take a type too;
- `max_work_group_size(kernel)` replaces `kernel_max_work_group_size`, and
  `max_num_groups` is only bounded by the global size;
- report sub-groups only for devices where kernels get a fixed sub-group width
  (through `cl_intel_required_subgroup_size`), and shuffles for the types the
  device supports with `cl_khr_subgroup_shuffle` (replacing `shfl_down_types`);
- conservative capabilities: float64 and unified memory per device, atomics;
- `kernel_function` keeps the backend value it was given;
- `copyto!` checks the lengths and returns the destination; forward
  `unsafe_free!`;
- cache the device properties that launches and queries need per device.

This needs SPIRVIntrinsics 1.1.3, whose 3-D builtins can be truncated without
producing vector types that SPIR-V doesn't have.
The platform is part of an `OpenCLBackend`'s configuration, but it was only
checked, with a warning, before allocating and launching, and `kernel_function`
replaced it by the active platform. Instead, work for a backend whose platform
isn't active now activates the platform's default device, as `KI.device!` would,
so that the arrays it creates can be used afterwards, and host queries answer for
that device. `get_backend` returns a backend for the array's platform.

Kernel queries answer for the device the kernel was compiled for, rather than for
the active one, and launching a kernel on another device throws an error instead
of failing in the driver. Also implement `KI.device(backend, A)`.
`KI.synchronize` has to let other tasks run while it waits, but `clFinish` blocks
the thread, and keeps the garbage collector from running on other threads. Like
CUDA.jl, `cl.finish(queue; blocking=false)` now enqueues a marker and busy-waits
for it briefly, then waits for it on a worker thread, in a GC-safe call, while the
task yields. The other uses of `cl.finish` keep blocking: they include finalizers,
which can't yield.
Julia doesn't specialize a method on `args...` that it only passes through, which
made every launch through KernelInterface's generic launch dispatch dynamically
(+1.2 µs and +1 kB per launch on CUDA).
@maleadt

maleadt commented Sep 30, 2026

Copy link
Copy Markdown
Member

#519

@maleadt maleadt closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants