Skip to content

Use GPUArrays 12's reductions, scans, sorting and findall - #665

Merged
maleadt merged 3 commits into
mainfrom
tb/gpuarrays-12
Oct 10, 2026
Merged

maleadt merged 3 commits into
mainfrom
tb/gpuarrays-12

Conversation

@maleadt

@maleadt maleadt commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Ports oneAPI.jl to GPUArrays 12. GPUArrays 12 implements Base's reductions, scans, sorting, findall and logical indexing once for every GPU array, on AcceleratedKernels 0.5 (JuliaGPU/GPUArrays.jl#790), and no longer calls the GPUArrays.mapreducedim! hook.

  • Deletes oneAPI's reduction kernel, its AcceleratedKernels wrappers for sorting and scans (written for AK 0.4's API), and findall/logical indexing. Drops the direct AcceleratedKernels dependency; its oneAPI extension still loads through GPUArrays.
  • The GPUArrays RNG is now constructed directly (RNG{oneArray}()), since GPUArrays 12 removes GPUArrays.default_rng and the RNG(state) constructor.

The Intel workarounds. oneAPI's own kernels worked around three problems (#563). AcceleratedKernels doesn't carry them:

  • Scans with a block size of 128 or more gave wrong results.
  • Writing 1- and 2-byte values to local memory clobbered adjacent bytes.
  • The Aurora LTS compiler miscompiled strided reads.

On an Iris Xe, AK's scans (n = 1 to 3M, block sizes 128/256, along dims) and sub-word reductions (Bool/Int8/UInt8/Int16/UInt16 with |, &, xor, max, min, several dims) are correct. The LTS problem can't be reproduced there. test/array.jl now has tests for the cases each workaround covered, so the Aurora run on this PR shows whether any of them is still needed. If one is, it goes into AcceleratedKernels.

Tested on an Iris Xe (Julia 1.13), full suite: 13,393 pass, 4 fail, 3 errors, all fixed elsewhere:

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Your PR requires formatting changes to meet the project's style guidelines.
Please consider running Runic (git runic main) to apply these changes.

Click here to view the suggested changes.
diff --git a/src/random.jl b/src/random.jl
index e183773..1f28168 100644
--- a/src/random.jl
+++ b/src/random.jl
@@ -4,9 +4,9 @@ using Random
 function gpuarrays_rng()
     dev = device()
     rngs = get!(task_local_storage(), :oneAPI_GLOBAL_RNGs) do
-        Dict{ZeDevice,GPUArrays.RNG{oneArray}}()
+        Dict{ZeDevice, GPUArrays.RNG{oneArray}}()
     end
-    get!(() -> GPUArrays.RNG{oneArray}(), rngs, dev)
+    return get!(() -> GPUArrays.RNG{oneArray}(), rngs, dev)
 end
 
 # GPUArrays in-place
diff --git a/test/array.jl b/test/array.jl
index c014cf2..0c43b71 100644
--- a/test/array.jl
+++ b/test/array.jl
@@ -88,41 +88,41 @@ end
 end
 
 @testset "aliasing" begin
-  x = oneArray([1, 2])
-  y = view(x, 2:2)
-  @test Base.mightalias(x, x)
-  @test Base.mightalias(x, y)
-  z = view(x, 1:1)
-  @test Base.mightalias(x, z)
-  @test !Base.mightalias(y, z)
-
-  a = copy(y)::typeof(x)
-  @test !Base.mightalias(x, a)
-  b = Base.unaliascopy(y)::typeof(y)
-  @test !Base.mightalias(x, b)
-
-  # contiguous views are oneArrays with an offset into the parent's memory,
-  # which should still alias wrapped arrays (like SubArrays) of that memory
-  x = oneArray(1:16)
-  @test Base.mightalias(view(x, 2:16), view(x, 15:-1:1))
-  @test Base.mightalias(view(x, 1:2:15), view(x, 2:9))
-  @test Base.mightalias(view(x, 2:16), view(reinterpret(Int32, x), 1:2:31))
-  @test !Base.mightalias(view(x, 2:16), view(oneArray(1:16), 15:-1:1))
-
-  # so in-place broadcasts between them should make a copy first
-  n = 2^20
-  x = oneArray{Float32}(1:n)
-  view(x, 2:n) .= view(x, n-1:-1:1)
-  @test Array(x) == [1; n-1:-1:1]
-
-  # empty arrays alias nothing, also on Julia 1.10
-  @test !Base.mightalias(oneArray(Int[]), oneArray(Float32[]))
-  @test !Base.mightalias(view(oneArray(zeros(Float32, 2, 0)), 1:1, :), oneArray(Int[]))
-
-  # disjoint parts of one array may be each other's source and destination
-  x = oneArray(collect(1:10))
-  @test Array(sum!(view(x, 1:1), view(x, 2:10))) == [54]
-  @test Array(cumsum!(view(x, 1:5), view(x, 6:10))) == cumsum(6:10)
+    x = oneArray([1, 2])
+    y = view(x, 2:2)
+    @test Base.mightalias(x, x)
+    @test Base.mightalias(x, y)
+    z = view(x, 1:1)
+    @test Base.mightalias(x, z)
+    @test !Base.mightalias(y, z)
+
+    a = copy(y)::typeof(x)
+    @test !Base.mightalias(x, a)
+    b = Base.unaliascopy(y)::typeof(y)
+    @test !Base.mightalias(x, b)
+
+    # contiguous views are oneArrays with an offset into the parent's memory,
+    # which should still alias wrapped arrays (like SubArrays) of that memory
+    x = oneArray(1:16)
+    @test Base.mightalias(view(x, 2:16), view(x, 15:-1:1))
+    @test Base.mightalias(view(x, 1:2:15), view(x, 2:9))
+    @test Base.mightalias(view(x, 2:16), view(reinterpret(Int32, x), 1:2:31))
+    @test !Base.mightalias(view(x, 2:16), view(oneArray(1:16), 15:-1:1))
+
+    # so in-place broadcasts between them should make a copy first
+    n = 2^20
+    x = oneArray{Float32}(1:n)
+    view(x, 2:n) .= view(x, (n - 1):-1:1)
+    @test Array(x) == [1; (n - 1):-1:1]
+
+    # empty arrays alias nothing, also on Julia 1.10
+    @test !Base.mightalias(oneArray(Int[]), oneArray(Float32[]))
+    @test !Base.mightalias(view(oneArray(zeros(Float32, 2, 0)), 1:1, :), oneArray(Int[]))
+
+    # disjoint parts of one array may be each other's source and destination
+    x = oneArray(collect(1:10))
+    @test Array(sum!(view(x, 1:1), view(x, 2:10))) == [54]
+    @test Array(cumsum!(view(x, 1:5), view(x, 6:10))) == cumsum(6:10)
 end
 
 @testset "shared buffers & unsafe_wrap" begin
@@ -222,7 +222,7 @@ end
     dA = oneArray(A)
     @test dA == transpose(oneArray(collect(transpose(A))))
     @test sum(transpose(dA)) == sum(A)
-    @test Array(sum(transpose(dA); dims=1)) == sum(transpose(A); dims=1)
+    @test Array(sum(transpose(dA); dims = 1)) == sum(transpose(A); dims = 1)
     for (sz, dts) in (
             ((2, 512, 64), ((1, 3), (2,), (3,), (2, 3), (1, 2, 3), 1)),
             ((3, 256, 48), ((1, 3), (2,), (1, 2, 3))),
@@ -261,7 +261,7 @@ end
     end
     for T in (Bool, Int8, Int16), (sz, dims) in (((64, 1000), 1), ((64, 1000), 2), ((17, 33, 65), (1, 3)))
         A = T == Bool ? rand(Bool, sz...) : rand(T(0):T(1), sz...)
-        @test Array(reduce(xor, oneArray(A); dims, init=zero(T))) == reduce(xor, A; dims, init=zero(T))
+        @test Array(reduce(xor, oneArray(A); dims, init = zero(T))) == reduce(xor, A; dims, init = zero(T))
         @test Array(maximum(oneArray(A); dims)) == maximum(A; dims)
     end
 end
@@ -273,6 +273,6 @@ end
         @test Array(cumsum(oneArray(A))) == cumsum(A)
     end
     A = rand(Int32(0):Int32(3), 300, 700)
-    @test Array(cumsum(oneArray(A); dims=1)) == cumsum(A; dims=1)
-    @test Array(cumsum(oneArray(A); dims=2)) == cumsum(A; dims=2)
+    @test Array(cumsum(oneArray(A); dims = 1)) == cumsum(A; dims = 1)
+    @test Array(cumsum(oneArray(A); dims = 2)) == cumsum(A; dims = 2)
 end

@codecov

codecov Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.26%. Comparing base (b0ace13) to head (1bb8861).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #665      +/-   ##
==========================================
- Coverage   81.14%   80.26%   -0.89%     
==========================================
  Files          57       51       -6     
  Lines        4122     3938     -184     
==========================================
- Hits         3345     3161     -184     
  Misses        777      777              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

GPUArrays 12 removes the default_rng interface and the RNG constructor that took
a state array; its RNG is now stateless and built as RNG{oneArray}().
GPUArrays 12 implements these for every GPU array on AcceleratedKernels and no
longer calls the mapreducedim! hook, so oneAPI's own reduction kernel, its
AcceleratedKernels wrappers and its findall/logical indexing go, and so does
the direct dependency on AcceleratedKernels (its oneAPI extension still loads
through GPUArrays).

The Intel workarounds those kernels carried (scan block size 64, widening
sub-word types in local memory, densifying strided inputs on the Aurora LTS
stack) are not needed by AcceleratedKernels on an Iris Xe; the tests now
check the cases they covered, so CI on Aurora shows whether they still are.
Base.dataids returned the array's own pointer, which missed overlaps between
a contiguous view (a oneArray at an offset) and a wrapped array of the same
memory, and the byte-range mightalias let empty arrays alias. GPUArrays 12.1
defines `Base.dataids` and `Base.mightalias` for every GPU array from where
its elements live, so implement that hook instead, reading the address from
the buffer.
@maleadt
maleadt merged commit e1e564d into main Oct 10, 2026
5 checks passed
@maleadt
maleadt deleted the tb/gpuarrays-12 branch October 10, 2026 06:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant