Skip to content

Adapt to LLVMDowngrader_jll 0.11 and GPUCompiler 2.8 - #1072

Merged
luraess merged 2 commits into
mainfrom
tb/llvm23-jlls
Sep 23, 2026
Merged

luraess merged 2 commits into
mainfrom
tb/llvm23-jlls

Conversation

@maleadt

@maleadt maleadt commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

Bumps the compat bounds to LLVMDowngrader_jll 0.11 (JuliaPackaging/Yggdrasil#14758, built against LLVM 23) and GPUCompiler 2.8 (JuliaGPU/GPUCompiler.jl#931). The downgrader C API is unchanged, so this is compat-only.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMDGPU.jl Benchmarks

Details
Benchmark suite Current: d1a2cf9 Previous: 4f0ab1a Ratio
amdgpu/synchronization/context/device 567.5 ns 562.5 ns 1.01
amdgpu/synchronization/stream/blocking 240 ns 235 ns 1.02
amdgpu/synchronization/stream/nonblocking 337.5 ns 320 ns 1.05
applications/bitonic_sort 1137386.25 ns 1098082.75 ns 1.04
applications/convolution 104711.25 ns 104229 ns 1.00
applications/floyd_warshall 8868676 ns 8878742.25 ns 1.00
applications/histogram 807982.25 ns 815134 ns 0.99
applications/prefix_sum 232332.75 ns 231900.75 ns 1.00
array/accumulate/Float32/1d 75291 ns 80493.75 ns 0.94
array/accumulate/Float32/dims=1 263556 ns 272491.25 ns 0.97
array/accumulate/Float32/dims=1L 79846 ns 88573.75 ns 0.90
array/accumulate/Float32/dims=2 70281 ns 81833.75 ns 0.86
array/accumulate/Float32/dims=2L 2751418.75 ns 2754656.25 ns 1.00
array/accumulate/Int64/1d 77256.25 ns 79688.75 ns 0.97
array/accumulate/Int64/dims=1 241468.25 ns 244706 ns 0.99
array/accumulate/Int64/dims=1L 83508.75 ns 84231.25 ns 0.99
array/accumulate/Int64/dims=2 83876.25 ns 86993.75 ns 0.96
array/accumulate/Int64/dims=2L 2892030.75 ns 3133859.25 ns 0.92
array/broadcast 71978.5 ns 71851 ns 1.00
array/construct 2115 ns 2385 ns 0.89
array/copy 37240.5 ns 38018 ns 0.98
array/copyto!/cpu_to_gpu 110716.5 ns 110996.5 ns 1.00
array/copyto!/gpu_to_cpu 120734 ns 111031.5 ns 1.09
array/copyto!/gpu_to_gpu 48610.75 ns 59370.75 ns 0.82
array/iteration/findall/bool 133906.75 ns 135379.5 ns 0.99
array/iteration/findall/int 140154.25 ns 150062.25 ns 0.93
array/iteration/findfirst/bool 181530 ns 152017.25 ns 1.19
array/iteration/findfirst/int 160627.5 ns 142527 ns 1.13
array/iteration/findmin/1d 113356.75 ns 106111.5 ns 1.07
array/iteration/findmin/2d 107201.5 ns 108821.5 ns 0.99
array/iteration/logical 240265.75 ns 243736 ns 0.99
array/iteration/scalar 292254.25 ns 300659.25 ns 0.97
array/permutedims/2d 70518.5 ns 71656 ns 0.98
array/permutedims/3d 70326 ns 71226 ns 0.99
array/permutedims/4d 72571 ns 73858.75 ns 0.98
array/random/rand/Float32 44275.5 ns 45950.75 ns 0.96
array/random/rand/Int64 52868.25 ns 54290.75 ns 0.97
array/random/rand!/Float32 64280.75 ns 65176 ns 0.99
array/random/rand!/Int64 70811 ns 66563.5 ns 1.06
array/random/randn/Float32 66636 ns 81206.25 ns 0.82
array/random/randn!/Float32 80228.75 ns 80368.5 ns 1.00
array/reductions/mapreduce/Float32/1d 79508.75 ns 95054 ns 0.84
array/reductions/mapreduce/Float32/dims=1 85838.75 ns 86063.75 ns 1.00
array/reductions/mapreduce/Float32/dims=1L 839962 ns 843989.5 ns 1.00
array/reductions/mapreduce/Float32/dims=2 81936.25 ns 75981 ns 1.08
array/reductions/mapreduce/Float32/dims=2L 139104.5 ns 138774.5 ns 1.00
array/reductions/mapreduce/Int64/1d 90881.5 ns 94306.5 ns 0.96
array/reductions/mapreduce/Int64/dims=1 84831.25 ns 86678.75 ns 0.98
array/reductions/mapreduce/Int64/dims=1L 844937 ns 846329.5 ns 1.00
array/reductions/mapreduce/Int64/dims=2 80881.25 ns 81706.25 ns 0.99
array/reductions/mapreduce/Int64/dims=2L 139254.5 ns 140432 ns 0.99
array/reductions/reduce/Float32/1d 82521.25 ns 94821.25 ns 0.87
array/reductions/reduce/Float32/dims=1 85478.75 ns 87331.25 ns 0.98
array/reductions/reduce/Float32/dims=1L 840351.75 ns 834664.5 ns 1.01
array/reductions/reduce/Float32/dims=2 82363.75 ns 76153.5 ns 1.08
array/reductions/reduce/Float32/dims=2L 138911.75 ns 139254.5 ns 1.00
array/reductions/reduce/Int64/1d 92781.25 ns 94776.5 ns 0.98
array/reductions/reduce/Int64/dims=1 84693.5 ns 86181 ns 0.98
array/reductions/reduce/Int64/dims=1L 843539.5 ns 839047 ns 1.01
array/reductions/reduce/Int64/dims=2 82026 ns 83326 ns 0.98
array/reductions/reduce/Int64/dims=2L 139549.5 ns 140177 ns 1.00
array/reverse/1d 43498 ns 42600.5 ns 1.02
array/reverse/1dL 56203.25 ns 72488.5 ns 0.78
array/reverse/1dL_inplace 78098.75 ns 78666 ns 0.99
array/reverse/1d_inplace 50593.25 ns 56155.75 ns 0.90
array/reverse/2d 48898.25 ns 49858.25 ns 0.98
array/reverse/2dL 81691 ns 82311.25 ns 0.99
array/reverse/2dL_inplace 89221.25 ns 89706.25 ns 0.99
array/reverse/2d_inplace 61108.5 ns 61980.75 ns 0.99
array/sorting/1d 333424.75 ns 332449.75 ns 1.00
gemm/tiled 1925307 ns 1904269.25 ns 1.01
gemm/tiled_unbounded 1928027.25 ns 1909711.75 ns 1.01
integration/byval/reference 39030 ns 39271 ns 0.99
integration/byval/slices=1 39330 ns 41411 ns 0.95
integration/byval/slices=2 135922 ns 132252 ns 1.03
integration/byval/slices=3 238752 ns 238713 ns 1.00
integration/volumerhs 4918610 ns 4997480 ns 0.98
kernel/indexing 55646 ns 28585.25 ns 1.95
kernel/indexing_checked 56660.75 ns 57423.5 ns 0.99
kernel/launch 1285 ns 1255 ns 1.02
kernel/rand 97281.5 ns 98873.75 ns 0.98
latency/import 1705064958 ns 1700382856 ns 1.00
latency/precompile 39927733725 ns 40071892794 ns 1.00
latency/ttfp 2307133138 ns 2305806765 ns 1.00
stencil/diffusion3d 1614670.25 ns 1623510.75 ns 0.99
stencil/diffusion3d_checked 1658811 ns 1662718.75 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@christiangnrd

Copy link
Copy Markdown
Member

Close #1080

@luraess

luraess commented Sep 22, 2026

Copy link
Copy Markdown
Member

Can you rebase on main to see that CI passes and we could merge

GPUCompiler is temporarily sourced from its tb/llvm23-jlls branch so that CI
can run against it before the release is tagged; drop the source once it is.
@luraess
luraess merged commit 9e8f14b into main Sep 23, 2026
5 of 6 checks passed
@luraess
luraess deleted the tb/llvm23-jlls branch September 23, 2026 19:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants