A kernel whose @Const argument is a ReshapedArray (e.g. vec(view(A, 1:4, 1:4))) fails to compile on CUDA; without @Const the same kernel works.
using CUDA, KernelAbstractions
@kernel function copy_const!(y, @Const(x))
i = @index(Global)
@inbounds y[i] = x[i]
end
@kernel function copy_plain!(y, x)
i = @index(Global)
@inbounds y[i] = x[i]
end
A = CuArray(Float32.(reshape(1:25, 5, 5)))
x = vec(view(A, 1:4, 1:4)) # a ReshapedArray of a SubArray
y = CuArray{Float32}(undef, 16)
copy_plain!(CUDABackend())(y, x; ndrange=16) # works
copy_const!(CUDABackend())(y, x; ndrange=16) # InvalidIRError
InvalidIRError: compiling MethodInstance for gpu_copy_const!(..., ::CuDeviceVector{Float32, 1}, ::Base.ReshapedArray{Float32, 1, SubArray{Float32, 2, CuDeviceMatrix{Float32, 1}, Tuple{UnitRange{Int64}, UnitRange{Int64}}, false}, Tuple{Base.MultiplicativeInverses.SignedMultiplicativeInverse{Int64}}}) resulted in invalid LLVM IR
Reason: unsupported call to a lazy-initialized function (call to ijl_rethrow)
Stacktrace:
[1] rethrow @ ./error.jl:71
...
[5] SignedMultiplicativeInverse @ ./multinverses.jl:54
...
[10] reshape @ ./reshapedarray.jl:128
[11] adapt_structure @ Adapt/src/wrappers.jl:18
[12] adapt @ Adapt/src/Adapt.jl:40
[13] adapt_structure @ Adapt/src/wrappers.jl:10
[14] adapt @ Adapt/src/Adapt.jl:40
[15] constify @ KernelAbstractions/src/KernelAbstractions.jl:464
KernelAbstractions 0.9.42 (and main's constify is the same), CUDA.jl 6.4.0, Julia 1.13.0, RTX 5080.
@Const(x) expands to x = constify(x) inside the kernel, and constify(x) = adapt(ConstAdaptor(), x). For a ReshapedArray, Adapt rebuilds the wrapper with reshape(adapt(to, parent(x)), size(x)), which recomputes the SignedMultiplicativeInverses; their constructor has an error path that builds a string, which cannot compile for the GPU. So any @Const argument that is a reshape (of a view, or of anything that is not a plain device array) fails. Possible fixes: rebuild the ReshapedArray without recomputing its inverses (e.g. Base.ReshapedArray(adapt(to, parent(x)), size(x), x.mi)) in an adapt_structure method for ConstAdaptor, or apply the const marking only to the innermost device array.
AcceleratedKernels hits this in reduce/count of such views, which GPUArrays now routes to it (for example GPUArrays' norm(view(x, ...), 0) tests on CuArray). Its rework (JuliaGPU/AcceleratedKernels.jl#133) works around it by applying @Const only to plain dense device arrays, and will drop that once this is fixed.
A kernel whose
@Constargument is aReshapedArray(e.g.vec(view(A, 1:4, 1:4))) fails to compile on CUDA; without@Constthe same kernel works.KernelAbstractions 0.9.42 (and
main'sconstifyis the same), CUDA.jl 6.4.0, Julia 1.13.0, RTX 5080.@Const(x)expands tox = constify(x)inside the kernel, andconstify(x) = adapt(ConstAdaptor(), x). For aReshapedArray, Adapt rebuilds the wrapper withreshape(adapt(to, parent(x)), size(x)), which recomputes theSignedMultiplicativeInverses; their constructor has an error path that builds a string, which cannot compile for the GPU. So any@Constargument that is a reshape (of a view, or of anything that is not a plain device array) fails. Possible fixes: rebuild theReshapedArraywithout recomputing its inverses (e.g.Base.ReshapedArray(adapt(to, parent(x)), size(x), x.mi)) in anadapt_structuremethod forConstAdaptor, or apply the const marking only to the innermost device array.AcceleratedKernels hits this in
reduce/countof such views, which GPUArrays now routes to it (for example GPUArrays'norm(view(x, ...), 0)tests on CuArray). Its rework (JuliaGPU/AcceleratedKernels.jl#133) works around it by applying@Constonly to plain dense device arrays, and will drop that once this is fixed.