Skip to content

Use GPUArrays 12's reductions, sorting, scans and findall - #1145

Open
maleadt wants to merge 2 commits into
mainfrom
tb/gpuarrays-12
Open

maleadt wants to merge 2 commits into
mainfrom
tb/gpuarrays-12

Conversation

@maleadt

@maleadt maleadt commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Ports AMDGPU.jl to GPUArrays 12, which implements Base's reductions, sorting, scans, findall, logical indexing and reverse once for every GPU array, on AcceleratedKernels 0.5 (JuliaGPU/GPUArrays.jl#790), and removes the GPUArrays.mapreducedim! hook.

  • Deletes AMDGPU's own reduction kernel, sort!/sortperm!, scans, findall, logical indexing and reverse (src/kernels/). The vector sort!/sortperm! were ambiguous with GPUArrays' methods, and the scans used AcceleratedKernels 0.4's API.
  • Drops the direct AcceleratedKernels dependency; its AMDGPU extension still loads through GPUArrays. (AMDGPU's current compat, 0.3.1, 0.4, couldn't load AK 0.5 anyway.)
  • GPUArrays.default_rng is gone in GPUArrays 12; the RNG lives in gpuarrays_rng().

Intended as a minor release.

Tested on gfx1036 (Julia 1.13): core/rocarray_base, hip_rocarray, the GPUArrays testsuite and the KernelAbstractions tests, 17,968 pass. The 3 cholesky(::Diagonal) errors are fixed in Adapt 4.7.3.

GPUArrays 12 removes default_rng, and nothing outside AMDGPU called AMDGPU's
method for it.
GPUArrays 12 implements these once for every GPU array on top of
AcceleratedKernels and removes the mapreducedim! hook, so delete AMDGPU's own
reduction kernel and its sort!, sortperm!, accumulate, findall, logical
indexing and reverse methods. The vector sort! and sortperm! methods were
ambiguous with GPUArrays', and the scans called AcceleratedKernels 0.4's API.
AMDGPU no longer depends on AcceleratedKernels directly; its AMDGPU extension
still loads through GPUArrays.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant