Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/src/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,4 +112,5 @@ KernelAbstractions.LinearLaunch
KernelAbstractions.NDLaunch
KernelAbstractions.select_launch
KernelAbstractions.launch_workgroupsize
KernelAbstractions.compiler_options
```
94 changes: 37 additions & 57 deletions docs/src/implementations.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,63 +78,43 @@ Adapt.adapt_storage(::CUDABackend, x) = adapt(CuArray, x)

## Launching `@kernel` kernels

A kernel written with [`@kernel`](@ref) receives a hidden context, a
`KernelAbstractions.CompilerMetadata` built by the backend's `mkcontext`, from which
[`@index`](@ref) computes its indices. By default (a context without a `launch`),
`@index` assumes that the kernel was launched on a 1-D grid of
`length(blocks(iterspace))` groups of `length(workitems(iterspace))` work-items. It then
decomposes the linear hardware ids into Cartesian positions, which takes integer divisions
when the `ndrange` is not known at compile time, and computes in `Int`.

A backend **may** launch kernels differently, and pass the `launch` keyword to the
`CompilerMetadata` constructor to tell `@index` how:

- [`NDLaunch{T}`](@ref KernelAbstractions.NDLaunch): the grid has the shape of the
iteration space (for as many dimensions as the backend's grid has, i.e. up to 3), so
`@index` doesn't need any divisions;
- [`LinearLaunch{T}`](@ref KernelAbstractions.LinearLaunch): the default 1-D grid.

Either way `@index` computes in `T`, e.g. `Int32`, which is faster on GPUs. The backend
has to implement the typed [`KI.get_group_id`](@ref KernelInterface.get_group_id) and
[`KI.get_local_id`](@ref KernelInterface.get_local_id) queries such that they compute in
`T` too, e.g. without checked conversions.

[`select_launch`](@ref KernelAbstractions.select_launch) chooses the launch from the
iteration space, whether the workgroup size will be tuned, and the limits of the backend
([`KI.max_work_group_size`](@ref KernelInterface.max_work_group_size),
[`KI.max_work_group_dims`](@ref KernelInterface.max_work_group_dims) and
[`KI.max_num_groups`](@ref KernelInterface.max_num_groups)). It doesn't depend on the
workgroup size a backend tunes afterwards, which keeps the context type (and thus the
compiled kernel) the same before and after tuning, as long as the backend tunes with
[`launch_workgroupsize`](@ref KernelAbstractions.launch_workgroupsize). A launch then
looks like this:

```julia
function (obj::KA.Kernel{MyBackend})(args...; ndrange = nothing, workgroupsize = nothing)
ndrange, workgroupsize, iterspace, dynamic = KA.launch_config(obj, ndrange, workgroupsize)
launch = KA.select_launch(obj, workgroupsize, iterspace)
ctx = KA.CompilerMetadata{KA.ndrange(obj), KA.DynamicCheck}(ndrange, iterspace; launch)
kernel = compile(obj.f, ctx, args...)

if KA.workgroupsize(obj) <: KA.DynamicSize && workgroupsize === nothing
threads = max_threads(kernel) # at most `KI.max_work_group_size(backend)`
workgroupsize = KA.launch_workgroupsize(backend, launch, threads, ndrange)
iterspace, dynamic = KA.partition(obj, ndrange, workgroupsize)
ctx = KA.CompilerMetadata{KA.ndrange(obj), KA.DynamicCheck}(ndrange, iterspace; launch)
end

groups, items = size(KA.blocks(iterspace)), size(KA.workitems(iterspace))
prod(groups) == 0 && return
if launch isa KA.NDLaunch
run(kernel, ctx, args...; groups, items) # padded to 3 dimensions
else
run(kernel, ctx, args...; groups = prod(groups), items = prod(items))
end
end
```

The POCL backend is an example. Backends that launch on an N-d grid **must not** override
`__validindex` or the `__index_*` functions, which dispatch on the launch.
KernelAbstractions launches [`@kernel`](@ref) kernels on any backend that implements
[KernelInterface](@ref kernelinterface): it partitions the `ndrange`, builds the kernel's
hidden context (a `KernelAbstractions.CompilerMetadata`), compiles the kernel with
[`KI.kernel_function`](@ref KernelInterface.kernel_function), tunes the workgroup size, and
launches it with [`KI.launch`](@ref KernelInterface.launch). A backend doesn't implement any
of that itself, but it needs KernelInterface's typed index queries and an N-d
[`KI.launch`](@ref KernelInterface.launch), and it must not override KernelAbstractions' index
functions (see below). It **may** customize the launch through:

- [`KI.launch_configuration`](@ref KernelInterface.launch_configuration): the workgroup size
used when the kernel has no static or given one. It receives the number of work-items in
the `ndrange` as `nitems`, e.g. to prefer more workgroups over larger ones.
- [`KernelAbstractions.compiler_options`](@ref): compiler options for a kernel, e.g. a
register hint derived from its static workgroup size.
- `KernelAbstractions.Scratchpad`, which backs [`@private`](@ref) arrays and has to be
implemented (`@device_override`) for a backend's device: e.g. a stack allocation, or a
`StaticArrays.MArray`.
- `Adapt.adapt_storage(::KernelAbstractions.ConstAdaptor, x)` for the backend's device
arrays, which implements [`@Const`](@ref).

[`@index`](@ref) computes its indices from how a kernel was launched: on a grid with the
shape of the iteration space ([`NDLaunch`](@ref KernelAbstractions.NDLaunch), for up to as
many dimensions as the backend's grid has), which doesn't need any divisions, or on a 1-D
grid ([`LinearLaunch`](@ref KernelAbstractions.LinearLaunch)). Either way it computes in a
narrow index type such as `Int32` when the iteration space fits, which is why the typed
[`KI.get_group_id`](@ref KernelInterface.get_group_id) and
[`KI.get_local_id`](@ref KernelInterface.get_local_id) queries have to compute in that type
too, as KernelInterface specifies. For the same reason, backends **must not** override
`__validindex` or the `__index_*` functions.

A backend can still implement `(obj::KernelAbstractions.Kernel{MyBackend})(args...; ndrange, workgroupsize)`
to launch kernels itself, e.g. while it is being ported to KernelInterface. That relies on
KernelAbstractions internals: it has to choose the launch with
[`select_launch`](@ref KernelAbstractions.select_launch), pass it to the kernel's context,
and tune the workgroup size with
[`launch_workgroupsize`](@ref KernelAbstractions.launch_workgroupsize), as the generic
launch in `src/backend_launch.jl` does.

Packages that customize the iteration space (with a custom `partition` and `expand`)
don't need to do anything for these launches: the index functions only compute the global
Expand Down
23 changes: 11 additions & 12 deletions src/KernelAbstractions.jl
Original file line number Diff line number Diff line change
Expand Up @@ -472,12 +472,8 @@ synchronize(backend)
Use [`workgroupsize`](@ref KernelAbstractions.workgroupsize), [`ndrange`](@ref KernelAbstractions.ndrange),
and [`backend`](@ref KernelAbstractions.backend) to inspect a kernel's static configuration.

!!! note
Backend implementations **must** implement:
```
(kernel::Kernel{<:NewBackend})(args...; ndrange=nothing, workgroupsize=nothing)
```
As well as the on-device functionality.
Kernels are launched on any backend that implements [KernelInterface](@ref kernelinterface);
see the [notes for backend implementations](@ref implementations_notes).
"""
struct Kernel{Backend, WorkgroupSize <: _Size, NDRange <: _Size, Fun}
backend::Backend
Expand Down Expand Up @@ -563,13 +559,18 @@ last (possibly partial) workgroup. Primarily used by backend implementations and
@assert ndrange !== nothing
blocks, workgroupsize, dynamic = NDIteration.partition(extents(ndrange), workgroupsize)

if static_ndrange <: StaticSize
# the number of blocks is only static if the workgroup size is too: a backend that tunes
# the workgroup size would otherwise change the type of the kernel's context
if static_ndrange <: StaticSize && static_workgroupsize <: StaticSize
static_blocks = StaticSize{blocks}
blocks = nothing
mapping = NDIteration.static_mapping(ndrange)
else
static_blocks = DynamicSize
blocks = CartesianIndices(blocks)
end
if static_ndrange <: StaticSize
mapping = NDIteration.static_mapping(ndrange)
else
mapping = NDIteration.dynamic_mapping(ndrange)
end

Expand Down Expand Up @@ -611,10 +612,6 @@ function __workitems_iterspace end
end
end

# for reflection
function mkcontext end
function launch_config end

include("macros.jl")
include("spawn.jl")

Expand Down Expand Up @@ -644,6 +641,8 @@ automatically when a kernel is launched.
argconvert(k::Kernel{T}, arg) where {T} =
error("Don't know how to convert arguments for Kernel{$T}")

include("backend_launch.jl")

# Enzyme support
supports_enzyme(::Backend) = false
function __fake_compiler_job end
Expand Down
127 changes: 127 additions & 0 deletions src/backend_launch.jl
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
###
# Launching `@kernel` kernels on a KernelInterface backend
#
# Every backend implementing KernelInterface launches `@kernel` kernels with the methods
# below. Backends customize them through the hooks documented in `implementations.md`
# (`compiler_options`, and KernelInterface's `launch_configuration`) instead of
# reimplementing the launch.
###

"""
mkcontext(kernel::Kernel, ndrange, iterspace, [launch])

The hidden context argument for launching `kernel` over `ndrange`, partitioned as
`iterspace`, with the launch configuration `launch` (see [`select_launch`](@ref)).
"""
mkcontext(kernel::Kernel, _ndrange, iterspace) =
CompilerMetadata{ndrange(kernel), DynamicCheck}(_ndrange, iterspace)
mkcontext(kernel::Kernel, _ndrange, iterspace, launch) =
CompilerMetadata{ndrange(kernel), DynamicCheck}(_ndrange, iterspace; launch)
mkcontext(kernel::Kernel, I, _ndrange, iterspace, ::Dynamic) where {Dynamic} =
CompilerMetadata{ndrange(kernel), Dynamic}(I, _ndrange, iterspace)

"""
launch_config(kernel::Kernel, ndrange, workgroupsize)

Normalize the launch arguments of `kernel`, and partition the `ndrange`. Returns the
`ndrange` (`nothing` if it's static), the `workgroupsize` (`nothing` if it will be tuned),
the iteration space and whether it needs bounds checks. If the workgroup size will be
tuned, the iteration space is preliminary: it uses the `ndrange` as the workgroup size.
"""
function launch_config(kernel::Kernel, _ndrange, _workgroupsize)
if _ndrange isa Integer
_ndrange = (_ndrange,)
end
if _workgroupsize isa Integer
_workgroupsize = (_workgroupsize,)
end

iterspace, dynamic = if workgroupsize(kernel) <: DynamicSize && _workgroupsize === nothing
# use the ndrange as preliminary workgroupsize for autotuning
partition(kernel, _ndrange, something(_ndrange, static_ndrange(kernel)))
else
# this also checks that a given ndrange agrees with a static one
partition(kernel, _ndrange, _workgroupsize)
end
if ndrange(kernel) <: StaticSize
_ndrange = nothing
end

return _ndrange, _workgroupsize, iterspace, dynamic
end

"""
compiler_options(kernel::Kernel)::NamedTuple

Backend-specific compiler options for compiling `kernel` with
[`KI.kernel_function`](@ref KernelInterface.kernel_function), e.g. a hint derived from its
static workgroup size (CUDA.jl passes `maxthreads`). Backends **may** implement this for
their backend type; the default is no options.
"""
compiler_options(::Kernel) = (;)

static_ndrange(kernel::Kernel) = ndrange(kernel) <: StaticSize ? get(ndrange(kernel)) : nothing

# the product of `dims`, saturated at `typemax(Int)`
function saturated_prod(dims::Dims)
n = 1
for d in dims
n, overflow = Base.mul_with_overflow(n, d)
overflow && return typemax(Int)
end
return n
end

argconvert(kernel::Kernel{<:KI.Backend}, arg) = KI.argconvert(backend(kernel), arg)

function (obj::Kernel{<:KI.Backend})(args::Vararg{Any, N}; ndrange = nothing, workgroupsize = nothing) where {N}
ndrange, workgroupsize, iterspace, dynamic = launch_config(obj, ndrange, workgroupsize)
# nothing to launch (or compile) for an empty ndrange
any(iszero, size(blocks(iterspace))) && return nothing

# launch on an N-d grid, computing indices in 32 bits, if possible. this doesn't depend
# on the tuned workgroup size, so the context (and thus the kernel) doesn't either.
launch = select_launch(obj, workgroupsize, iterspace)
if launch === NDLaunch{Int32}()
# the common case, specialized statically
launch_kernel(obj, NDLaunch{Int32}(), ndrange, workgroupsize, iterspace, args...)
else
launch_kernel(obj, launch, ndrange, workgroupsize, iterspace, args...)
end
return nothing
end

function launch_kernel(
obj::Kernel, launch, ndrange, _workgroupsize, iterspace, args::Vararg{Any, N}
) where {N}
b = backend(obj)

# this might not be the final context, since we may tune the workgroupsize
ctx = mkcontext(obj, ndrange, iterspace, launch)
kernel = compile(obj, ctx, args...)

# tune the workgroup size, keeping the context type (and thus the kernel) the same
if workgroupsize(obj) <: DynamicSize && _workgroupsize === nothing
range = something(ndrange, static_ndrange(obj))
threads = KI.launch_configuration(kernel; nitems = saturated_prod(extents(range))).workgroupsize
iterspace, _ = partition(obj, ndrange, launch_workgroupsize(b, launch, threads, range))
ctx = mkcontext(obj, ndrange, iterspace, launch)
end

# launching through the `KI.Kernel` validates the sizes against the kernel's limits
groups = size(blocks(iterspace))
items = size(workitems(iterspace))
if launch isa NDLaunch
kernel(ctx, args...; numgroups = groups, workgroupsize = items)
else
kernel(ctx, args...; numgroups = prod(groups), workgroupsize = prod(items))
end
return nothing
end

@inline function compile(obj::Kernel, ctx, args::Vararg{Any, N}) where {N}
b = backend(obj)
f = KI.argconvert(b, obj.f)
tt = Tuple{Core.Typeof(KI.argconvert(b, ctx)), map(arg -> Core.Typeof(KI.argconvert(b, arg)), args)...}
return KI.kernel_function(b, f, tt; compiler_options(obj)...)
end
86 changes: 0 additions & 86 deletions src/pocl/backend.jl
Original file line number Diff line number Diff line change
Expand Up @@ -130,87 +130,6 @@ KI.supports_atomics(::POCLBackend) = true

## Kernel Launch

function KA.mkcontext(kernel::KA.Kernel{POCLBackend}, _ndrange, iterspace)
return KA.CompilerMetadata{KA.ndrange(kernel), KA.DynamicCheck}(_ndrange, iterspace)
end
function KA.mkcontext(kernel::KA.Kernel{POCLBackend}, _ndrange, iterspace, launch)
return KA.CompilerMetadata{KA.ndrange(kernel), KA.DynamicCheck}(_ndrange, iterspace; launch)
end
function KA.mkcontext(
kernel::KA.Kernel{POCLBackend}, I, _ndrange, iterspace,
::Dynamic
) where {Dynamic}
return KA.CompilerMetadata{KA.ndrange(kernel), Dynamic}(I, _ndrange, iterspace)
end

function KA.launch_config(kernel::KA.Kernel{POCLBackend}, ndrange, workgroupsize)
if ndrange isa Integer
ndrange = (ndrange,)
end
if workgroupsize isa Integer
workgroupsize = (workgroupsize,)
end

# partition checked that the ndrange's agreed
if KA.ndrange(kernel) <: KA.StaticSize
ndrange = nothing
end

iterspace, dynamic = if KA.workgroupsize(kernel) <: KA.DynamicSize &&
workgroupsize === nothing
# use ndrange as preliminary workgroupsize for autotuning
KA.partition(kernel, ndrange, ndrange)
else
KA.partition(kernel, ndrange, workgroupsize)
end

return ndrange, workgroupsize, iterspace, dynamic
end

function (obj::KA.Kernel{POCLBackend})(args::Vararg{Any, N}; ndrange = nothing, workgroupsize = nothing) where {N}
ndrange, workgroupsize, iterspace, dynamic =
KA.launch_config(obj, ndrange, workgroupsize)
# the launch doesn't depend on the tuned workgroup size, so neither does the context
launch = KA.select_launch(obj, workgroupsize, iterspace)
launch_kernel(obj, launch, ndrange, workgroupsize, iterspace, args...)
return nothing
end

function launch_kernel(obj, launch, ndrange, workgroupsize, iterspace, args::Vararg{Any, N}) where {N}
# this might not be the final context, since we may tune the workgroupsize
ctx = KA.mkcontext(obj, ndrange, iterspace, launch)
kernel = @opencl launch = false obj.f(ctx, args...)

# figure out the optimal workgroupsize automatically
if KA.workgroupsize(obj) <: KA.DynamicSize && workgroupsize === nothing
wg_info = cl.work_group_info(kernel.fun, device())
wg_size_nd = KA.launch_workgroupsize(KA.backend(obj), launch, wg_info.size, ndrange)
iterspace, dynamic = KA.partition(obj, ndrange, wg_size_nd)
ctx = KA.mkcontext(obj, ndrange, iterspace, launch)
end

groups = size(KA.blocks(iterspace))
items = size(KA.workitems(iterspace))
if prod(groups) == 0
return nothing
end

# Launch kernel
if launch isa KA.NDLaunch
local_size = pad3(items)
global_size = local_size .* pad3(groups)
else
local_size = prod(items)
global_size = prod(groups) * local_size
end
event = kernel(ctx, args...; global_size, local_size)
wait(event)
cl.clReleaseEvent(event)
return nothing
end

pad3(t::Tuple) = (t..., ntuple(_ -> 1, 3 - length(t))...)

KI.argconvert(::POCLBackend, arg) = clconvert(arg)

function KI.kernel_function(backend::POCLBackend, f::F, tt::TT = Tuple{}; name = nothing, kwargs...) where {F, TT}
Expand Down Expand Up @@ -343,9 +262,4 @@ end
POCL._print(args...)
end


## Other

KA.argconvert(::KA.Kernel{POCLBackend}, arg) = clconvert(arg)

end
2 changes: 1 addition & 1 deletion test/codegen_checks.jl
Original file line number Diff line number Diff line change
Expand Up @@ -174,7 +174,7 @@ end
@test @filecheck begin
@check "define spir_kernel void @{{.*}}gpu_codegen_global_linear"
@check "udiv i32"
@device_code_llvm debuginfo = :none KernelAbstractions.POCL.POCLKernels.launch_kernel(
@device_code_llvm debuginfo = :none KernelAbstractions.launch_kernel(
kernel, KernelAbstractions.LinearLaunch{Int32}(), ndrange, workgroupsize, iterspace, B
)
KernelAbstractions.synchronize(backend)
Expand Down
Loading
Loading