Skip to content

Wrong results for overflowing Int16 arithmetic in a reduction kernel #670

Description

@maleadt

On an Iris Xe (i7-1165G7), a reduction of abs2 over Complex{Int16} values gives wrong results when abs2 overflows Int16, in AcceleratedKernels' generic mapreducedim! kernel (the one used for sources without strides). Julia's IR is correct (wrapping mul i16/add i16 feeding llvm.smax.i16), so this looks like a miscompile further down, in the SPIR-V translation or IGC.

GPUArrays 12's testsuite hit it (reductions/mapreducedim!, Base.mapreducedim!(abs2, max, R, adjoint(A)) with Complex{Int16}). GPUArrays 12.0.1 keeps those values small.

using oneAPI
function check(label, gen, f)
    bad = 0
    for _ in 1:50
        A = gen()
        R0 = zeros(Int16, 1, 3)
        mk = A -> view(permutedims(A), [1, 2, 3, 4], :)   # not strided: AK's generic kernel
        c = Base.mapreducedim!(f, max, copy(R0), mk(A))
        g = Array(Base.mapreducedim!(f, max, oneArray(R0), mk(oneArray(A))))
        c == g || (bad += 1)
    end
    println(rpad(label, 30), bad, "/50")
end
function main()
    small = () -> complex.(rand(Int16(-100):Int16(100), 3, 4), rand(Int16(-100):Int16(100), 3, 4))
    full = () -> rand(Complex{Int16}, 3, 4)
    check("small abs2", small, abs2)
    check("full abs2", full, abs2)
    check("full re*re", full, x -> real(x) * real(x))
    check("full re*re+im*im", full, x -> real(x) * real(x) + imag(x) * imag(x))
    check("full re*im", full, x -> real(x) * imag(x))
end
main()
small abs2                    0/50
full abs2                     50/50
full re*re                    0/50
full re*re+im*im              49/50
full re*im                    0/50

The kernel's loop body, from @device_code_llvm:

  %.unpack = load i16, ptr addrspace(1) %16, align 2
  %.unpack73 = load i16, ptr addrspace(1) %.elt72, align 2
  %17 = mul i16 %.unpack, %.unpack
  %18 = mul i16 %.unpack73, %.unpack73
  %19 = add i16 %18, %17
  %20 = call i16 @llvm.smax.i16(i16 %19, i16 %value_phi24)

Hand-written @oneapi kernels with the same arithmetic (max over abs2 of Complex{Int16} loaded from a 2D array, with the seed passed as an argument or as a constant) compute the right result, so I couldn't reduce it below AK's kernel yet.

oneAPI 2.9.2 and the GPUArrays 12 port (#665), AcceleratedKernels 0.5.1, Julia 1.12.7 and 1.13.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions