On an Iris Xe (i7-1165G7), a reduction of abs2 over Complex{Int16} values gives wrong results when abs2 overflows Int16, in AcceleratedKernels' generic mapreducedim! kernel (the one used for sources without strides). Julia's IR is correct (wrapping mul i16/add i16 feeding llvm.smax.i16), so this looks like a miscompile further down, in the SPIR-V translation or IGC.
GPUArrays 12's testsuite hit it (reductions/mapreducedim!, Base.mapreducedim!(abs2, max, R, adjoint(A)) with Complex{Int16}). GPUArrays 12.0.1 keeps those values small.
using oneAPI
function check(label, gen, f)
bad = 0
for _ in 1:50
A = gen()
R0 = zeros(Int16, 1, 3)
mk = A -> view(permutedims(A), [1, 2, 3, 4], :) # not strided: AK's generic kernel
c = Base.mapreducedim!(f, max, copy(R0), mk(A))
g = Array(Base.mapreducedim!(f, max, oneArray(R0), mk(oneArray(A))))
c == g || (bad += 1)
end
println(rpad(label, 30), bad, "/50")
end
function main()
small = () -> complex.(rand(Int16(-100):Int16(100), 3, 4), rand(Int16(-100):Int16(100), 3, 4))
full = () -> rand(Complex{Int16}, 3, 4)
check("small abs2", small, abs2)
check("full abs2", full, abs2)
check("full re*re", full, x -> real(x) * real(x))
check("full re*re+im*im", full, x -> real(x) * real(x) + imag(x) * imag(x))
check("full re*im", full, x -> real(x) * imag(x))
end
main()
small abs2 0/50
full abs2 50/50
full re*re 0/50
full re*re+im*im 49/50
full re*im 0/50
The kernel's loop body, from @device_code_llvm:
%.unpack = load i16, ptr addrspace(1) %16, align 2
%.unpack73 = load i16, ptr addrspace(1) %.elt72, align 2
%17 = mul i16 %.unpack, %.unpack
%18 = mul i16 %.unpack73, %.unpack73
%19 = add i16 %18, %17
%20 = call i16 @llvm.smax.i16(i16 %19, i16 %value_phi24)
Hand-written @oneapi kernels with the same arithmetic (max over abs2 of Complex{Int16} loaded from a 2D array, with the seed passed as an argument or as a constant) compute the right result, so I couldn't reduce it below AK's kernel yet.
oneAPI 2.9.2 and the GPUArrays 12 port (#665), AcceleratedKernels 0.5.1, Julia 1.12.7 and 1.13.
On an Iris Xe (i7-1165G7), a reduction of
abs2overComplex{Int16}values gives wrong results whenabs2overflows Int16, in AcceleratedKernels' genericmapreducedim!kernel (the one used for sources without strides). Julia's IR is correct (wrappingmul i16/add i16feedingllvm.smax.i16), so this looks like a miscompile further down, in the SPIR-V translation or IGC.GPUArrays 12's testsuite hit it (
reductions/mapreducedim!,Base.mapreducedim!(abs2, max, R, adjoint(A))withComplex{Int16}). GPUArrays 12.0.1 keeps those values small.The kernel's loop body, from
@device_code_llvm:Hand-written
@oneapikernels with the same arithmetic (maxoverabs2ofComplex{Int16}loaded from a 2D array, with the seed passed as an argument or as a constant) compute the right result, so I couldn't reduce it below AK's kernel yet.oneAPI 2.9.2 and the GPUArrays 12 port (#665), AcceleratedKernels 0.5.1, Julia 1.12.7 and 1.13.