Skip to content

Kernels fail to compile for sources with a bits-union element type (Union{Missing, Bool}) #134

Description

@maleadt

Every AK kernel that marks its source @Const fails to compile on CUDA when the source's element
type is a bits union such as Union{Missing, Bool}. That covers the whole-array reduction, all
four launch shapes of the reductions along dims, any/all with ConcurrentWrite, and merge
sort.

using CUDA
import AcceleratedKernels as AK

x = CuArray(rand([true, false, missing], 1000, 300))
code(x) = x === missing ? 0x01 : x ? 0x02 : 0x00      # any bits-type map

AK.mapreduce(code, max, x; init=0x00, neutral=0x00)          # InvalidIRError (_mapreduce_block!)
AK.mapreduce(code, max, x; dims=1, init=0x00, neutral=0x00)  # InvalidIRError (_mapreduce_nd_by_block!)
AK.mapreduce(code, max, x; dims=2, init=0x00, neutral=0x00)  # InvalidIRError (_mapreduce_nd_by_thread!)
AK.any(ismissing, x)                                          # InvalidIRError (_any_global!)
AK.merge_sort!(CuArray(repeat(Union{Missing, Int32}[missing, 1], 600)))   # InvalidIRError
InvalidIRError: compiling MethodInstance for AcceleratedKernels.gpu__mapreduce_block!(...)
Reason: unsupported dynamic function invocation (call to pointerref(ptr::Core.LLVMPtr{T, A}, i::I, ::Val{align}) where {T, A, I, align} @ LLVM.Interop none:0)
Stacktrace:
 [1] unsafe_load            @ LLVM/src/interop/pointer.jl:87
 [2] unsafe_cached_load     @ CUDACore/src/device/pointer.jl:114
 [3] const_arrayref         @ CUDACore/src/device/array.jl:167
 [4] getindex               @ CUDACore/src/device/array.jl:216
 [5] macro expansion        @ AcceleratedKernels/src/reduce/mapreduce_1d_gpu.jl:24

Seen on AK main (5075d4e, 0.4.3) with CUDA.jl 6.4.0 (main), KernelAbstractions 0.9.42, Julia
1.13.0 and an RTX 5080. findall(ismissing, x) and the generic (non-strided) reduction kernel
work, since they read the source without @Const.

Cause. @Const(src) makes KernelAbstractions wrap the device array in
Base.Experimental.Const, and CUDA.jl's getindex on that is a cached load
(const_arrayref → unsafe_cached_load), which exists only for bits types: JuliaGPU/CUDA.jl#3290.
The kernels are in src/reduce/mapreduce_1d_gpu.jl (_mapreduce_block!), src/reduce/mapreduce_nd.jl
(_mapreduce_nd_by_thread_tiled_strided!, _mapreduce_nd_by_thread!, _mapreduce_nd_by_block!,
_mapreduce_nd_multigroup!), src/predicates.jl (_any_global!) and src/sort/merge_sort.jl,
src/sort/merge_sort_by_key.jl (the global merge kernels).

Why it matters now. GPUArrays is moving its reductions onto AK (JuliaGPU/GPUArrays.jl#790):
any(CuArray{Union{Missing, Bool}}) and all(...) then go through these kernels. They passed in
#790's first version only because CUDA.jl's own mapreducedim! served them.

Status. The AK rework (#133) works around it, marked WORKAROUND(CUDA.jl): the kernels apply
@Const only to bits-type inputs (src = isbitstype(eltype(src_arg)) ? KernelAbstractions.constify(src_arg) : src_arg,
which is what @Const expands to), with tests on every kernel shape. The workaround goes once
CUDA.jl#3290 is fixed; main (0.4) stays affected until then or until the rework is released.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions