GPUArrays implements findmin and findmax, but not their in-place forms, so findmin!(rval, rind, A) falls back to Base's findminmax! loop, which indexes scalars:
using JLArrays # or CUDA, etc.
A = JLArray(Float32[3 1; 2 4])
rval = JLArray(zeros(Float32, 1, 2))
rind = JLArray(fill(CartesianIndex(1, 1), 1, 2))
findmin!(rval, rind, A)
# ERROR: Scalar indexing is disallowed.
# Invocation of getindex resulted in scalar indexing of a GPU array.
(GPUArrays main, and with the AcceleratedKernels-based reductions of #790; Julia 1.13.)
An implementation must reproduce Base's init keyword, whose init=false does not fold the old values in: Base resets the index destination (fill!(rind, zero(eltype(keys(A))))), and "the initial values of Rval are not used if the corresponding indices in Rind are 0", so the old rval is ignored whenever the slice is not empty:
julia> findmin!([-10], [42], [3, 2]; init=false)
([2], [2])
julia> findmin!([0.5 5.0], [CartesianIndex(9, 9) CartesianIndex(9, 9)], [3.0 1; 2 4]; init=false)
([2.0 1.0], CartesianIndex{2}[CartesianIndex(2, 1) CartesianIndex(1, 2)])
With init=true, rval starts from the first slice and rind from zero, which gives the same result; either way the result is the reduction of A alone, and an empty A leaves rval and rind unchanged (Base returns early).
Two ways to implement it:
- Reduce
(f(x), i) pairs into a temporary array of tuples with AK.mapreducedim! in overwrite mode, as findmin(A; dims) does, then unzip it into rval and rind (rind through keys(A)). This needs no change in AcceleratedKernels, and costs one extra pass over the (small) output and a temporary of its size.
- Let
AK.mapreducedim! write a destination that is a pair of arrays, which saves that pass but adds a new destination contract to AcceleratedKernels.
Option 1 is the natural first step; findmin!/findmax! along dims are rarely performance-critical.
GPUArrays implements
findminandfindmax, but not their in-place forms, sofindmin!(rval, rind, A)falls back to Base'sfindminmax!loop, which indexes scalars:(GPUArrays
main, and with the AcceleratedKernels-based reductions of #790; Julia 1.13.)An implementation must reproduce Base's
initkeyword, whoseinit=falsedoes not fold the old values in: Base resets the index destination (fill!(rind, zero(eltype(keys(A))))), and "the initial values ofRvalare not used if the corresponding indices inRindare 0", so the oldrvalis ignored whenever the slice is not empty:With
init=true,rvalstarts from the first slice andrindfrom zero, which gives the same result; either way the result is the reduction ofAalone, and an emptyAleavesrvalandrindunchanged (Base returns early).Two ways to implement it:
(f(x), i)pairs into a temporary array of tuples withAK.mapreducedim!in overwrite mode, asfindmin(A; dims)does, then unzip it intorvalandrind(rindthroughkeys(A)). This needs no change in AcceleratedKernels, and costs one extra pass over the (small) output and a temporary of its size.AK.mapreducedim!write a destination that is a pair of arrays, which saves that pass but adds a new destination contract to AcceleratedKernels.Option 1 is the natural first step;
findmin!/findmax!alongdimsare rarely performance-critical.