Every AK kernel that marks its source @Const fails to compile on CUDA when the source's element
type is a bits union such as Union{Missing, Bool}. That covers the whole-array reduction, all
four launch shapes of the reductions along dims, any/all with ConcurrentWrite, and merge
sort.
using CUDA
import AcceleratedKernels as AK
x = CuArray(rand([true, false, missing], 1000, 300))
code(x) = x === missing ? 0x01 : x ? 0x02 : 0x00 # any bits-type map
AK.mapreduce(code, max, x; init=0x00, neutral=0x00) # InvalidIRError (_mapreduce_block!)
AK.mapreduce(code, max, x; dims=1, init=0x00, neutral=0x00) # InvalidIRError (_mapreduce_nd_by_block!)
AK.mapreduce(code, max, x; dims=2, init=0x00, neutral=0x00) # InvalidIRError (_mapreduce_nd_by_thread!)
AK.any(ismissing, x) # InvalidIRError (_any_global!)
AK.merge_sort!(CuArray(repeat(Union{Missing, Int32}[missing, 1], 600))) # InvalidIRError
InvalidIRError: compiling MethodInstance for AcceleratedKernels.gpu__mapreduce_block!(...)
Reason: unsupported dynamic function invocation (call to pointerref(ptr::Core.LLVMPtr{T, A}, i::I, ::Val{align}) where {T, A, I, align} @ LLVM.Interop none:0)
Stacktrace:
[1] unsafe_load @ LLVM/src/interop/pointer.jl:87
[2] unsafe_cached_load @ CUDACore/src/device/pointer.jl:114
[3] const_arrayref @ CUDACore/src/device/array.jl:167
[4] getindex @ CUDACore/src/device/array.jl:216
[5] macro expansion @ AcceleratedKernels/src/reduce/mapreduce_1d_gpu.jl:24
Seen on AK main (5075d4e, 0.4.3) with CUDA.jl 6.4.0 (main), KernelAbstractions 0.9.42, Julia
1.13.0 and an RTX 5080. findall(ismissing, x) and the generic (non-strided) reduction kernel
work, since they read the source without @Const.
Cause. @Const(src) makes KernelAbstractions wrap the device array in
Base.Experimental.Const, and CUDA.jl's getindex on that is a cached load
(const_arrayref → unsafe_cached_load), which exists only for bits types: JuliaGPU/CUDA.jl#3290.
The kernels are in src/reduce/mapreduce_1d_gpu.jl (_mapreduce_block!), src/reduce/mapreduce_nd.jl
(_mapreduce_nd_by_thread_tiled_strided!, _mapreduce_nd_by_thread!, _mapreduce_nd_by_block!,
_mapreduce_nd_multigroup!), src/predicates.jl (_any_global!) and src/sort/merge_sort.jl,
src/sort/merge_sort_by_key.jl (the global merge kernels).
Why it matters now. GPUArrays is moving its reductions onto AK (JuliaGPU/GPUArrays.jl#790):
any(CuArray{Union{Missing, Bool}}) and all(...) then go through these kernels. They passed in
#790's first version only because CUDA.jl's own mapreducedim! served them.
Status. The AK rework (#133) works around it, marked WORKAROUND(CUDA.jl): the kernels apply
@Const only to bits-type inputs (src = isbitstype(eltype(src_arg)) ? KernelAbstractions.constify(src_arg) : src_arg,
which is what @Const expands to), with tests on every kernel shape. The workaround goes once
CUDA.jl#3290 is fixed; main (0.4) stays affected until then or until the rework is released.
Every AK kernel that marks its source
@Constfails to compile on CUDA when the source's elementtype is a bits union such as
Union{Missing, Bool}. That covers the whole-array reduction, allfour launch shapes of the reductions along
dims,any/allwithConcurrentWrite, and mergesort.
Seen on AK
main(5075d4e, 0.4.3) with CUDA.jl 6.4.0 (main), KernelAbstractions 0.9.42, Julia1.13.0 and an RTX 5080.
findall(ismissing, x)and the generic (non-strided) reduction kernelwork, since they read the source without
@Const.Cause.
@Const(src)makes KernelAbstractions wrap the device array inBase.Experimental.Const, and CUDA.jl'sgetindexon that is a cached load(
const_arrayref→unsafe_cached_load), which exists only for bits types: JuliaGPU/CUDA.jl#3290.The kernels are in
src/reduce/mapreduce_1d_gpu.jl(_mapreduce_block!),src/reduce/mapreduce_nd.jl(
_mapreduce_nd_by_thread_tiled_strided!,_mapreduce_nd_by_thread!,_mapreduce_nd_by_block!,_mapreduce_nd_multigroup!),src/predicates.jl(_any_global!) andsrc/sort/merge_sort.jl,src/sort/merge_sort_by_key.jl(the global merge kernels).Why it matters now. GPUArrays is moving its reductions onto AK (JuliaGPU/GPUArrays.jl#790):
any(CuArray{Union{Missing, Bool}})andall(...)then go through these kernels. They passed in#790's first version only because CUDA.jl's own
mapreducedim!served them.Status. The AK rework (#133) works around it, marked
WORKAROUND(CUDA.jl): the kernels apply@Constonly to bits-type inputs (src = isbitstype(eltype(src_arg)) ? KernelAbstractions.constify(src_arg) : src_arg,which is what
@Constexpands to), with tests on every kernel shape. The workaround goes onceCUDA.jl#3290 is fixed;
main(0.4) stays affected until then or until the rework is released.