Skip to content

KernelInterface 0.4: a better API for writing kernels #810

Description

@maleadt

KernelInterface 0.3 (#800) is mostly about the contract between KernelAbstractions and the back-ends. KI now validates and sizes launches itself, so back-ends implement KI.launch with checked 3-D sizes instead of the whole launch path. The semantics that differed between back-ends are written down (work-group limits, typed index queries, sub-groups, capability defaults), and a testsuite checks them. On top of that, #801 lets KA launch @kernel kernels on any KI back-end, which shrinks CUDA's KA extension from 127 to 37 lines.

KI is also meant for users who want to write kernels without KA's macros, with more control or less overhead. 0.3 doesn't serve them as well yet. Some of the fixes are breaking, but since we know of no direct KI users outside the back-ends yet, I propose releasing KI 0.4 in a couple of weeks. It should be out before KA 0.10 is registered, so that KA 0.10 only ever targets 0.4.

Proposed for 0.4

  1. ndrange has different semantics in KA and KI. KA masks work-items outside the ndrange by default. KI rounds it up to whole work-groups and leaves bounds checks to the kernel. That is the right choice at this level: the kernel can guard its memory accesses while all work-items still reach barriers. But without an explicit bounds check, a kernel ported from KA and launched with KI.@launch ndrange=... can access memory out of bounds when the range is padded. Rename the keyword, or at least make the difference hard to miss in the docs.

  2. No private-memory primitive in KI. KI has localmemory, but per-work-item scratch storage is only exposed through KA.Scratchpad(ctx, T, Val(dims)), which takes KA's hidden context. Back-ends override it and ignore the context (CUDA returns an MArray, PoCL a stack allocation). Add a portable KI primitive without the context, with defined lifetime and initialization, and lower @private to it. Back-ends then no longer need a KA hook for it.

  3. @Const has no KI equivalent. @Const promises that the memory an argument refers to is neither written by the kernel nor aliased by other arguments. Back-ends implement it with an Adapt rule for KA.ConstAdaptor. KI users could use the same, e.g. to get ld.global.nc loads on CUDA. This could become a KI primitive with the same promise, leaving the annotation and the argument traversal in KA. It is less clear-cut than private memory, so it could also wait.

  4. supports_subgroups implies a fixed width. It covers the sub-group queries and sub_group_barrier, and also promises that every compiled kernel uses the width reported by sub_group_size(backend). A back-end that can't guarantee the width has to report no sub-group support at all, which is what happens on OpenCL with NVIDIA. Kernels that read the width at run time don't need that guarantee. Split it into support for sub-group operations and a separate guarantee of a fixed width.

  5. Launch keywords. Keywords that KI doesn't know go to KI.launch on a Kernel call, but to kernel_function on KI.@launch. So back-end launch options such as CUDA's stream need launch=false and a second call. Adding a KI keyword later also changes what happens to a back-end keyword of the same name: KernelInterface 0.3: a tighter back-end contract #800 proposes dynamic local memory as a Kernel keyword, and calling it shmem would take over a keyword that is currently forwarded to CUDA. Either separate compiler and launch options explicitly (e.g. launch_options=(; stream)), or add a rule that new KI keywords only come in breaking releases.

Non-breaking (can go in any release)

  • An example that uses KI without KA. Both KI examples load KA, and examples/histogram.jl takes its atomics from KA.
  • Make clear that supports_atomics does not guarantee local-memory atomics, even though the histogram example uses them.
  • A guide for kernel authors that brings the existing contracts together: for each function, who calls it and who implements it, and which behavior is unspecified (shuffles from lanes that don't exist, how work-items map onto sub-groups). It should also point out that the index queries default to Int. Since Launch @kernel kernels on N-d grids, indexing in 32 bits #797, KA picks Int32 by itself when the indices fit, so a KI port of a KA kernel can end up with 64-bit index arithmetic where KA used 32-bit.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions