Adapt to LLVMDowngrader_jll 0.11 and GPUCompiler 2.8 - #965
Merged
Merged
Conversation
GPUCompiler is temporarily sourced from its tb/llvm23-jlls branch so that CI can run against it before the release is tagged; drop the source once it is.
Contributor
There was a problem hiding this comment.
Metal Benchmarks
Details
| Benchmark suite | Current: f498a72 | Previous: 8b6fa21 | Ratio |
|---|---|---|---|
array/accumulate/Float32/1d |
405083 ns |
408042 ns |
0.99 |
array/accumulate/Float32/dims=1 |
378458 ns |
375167 ns |
1.01 |
array/accumulate/Float32/dims=1L |
8802625 ns |
8807041 ns |
1.00 |
array/accumulate/Float32/dims=2 |
441167 ns |
441625 ns |
1.00 |
array/accumulate/Float32/dims=2L |
2595750 ns |
2593291 ns |
1.00 |
array/accumulate/Int64/1d |
839875 ns |
842500 ns |
1.00 |
array/accumulate/Int64/dims=1 |
937709 ns |
708375 ns |
1.32 |
array/accumulate/Int64/dims=1L |
9517542 ns |
9500083 ns |
1.00 |
array/accumulate/Int64/dims=2 |
1240708 ns |
1243500 ns |
1.00 |
array/accumulate/Int64/dims=2L |
6531708 ns |
6527833 ns |
1.00 |
array/broadcast |
223167 ns |
231250 ns |
0.97 |
array/construct |
2250 ns |
2208 ns |
1.02 |
array/permutedims/2d |
400375 ns |
398708 ns |
1.00 |
array/permutedims/3d |
1018750 ns |
1015166 ns |
1.00 |
array/permutedims/4d |
1154125 ns |
1206375 ns |
0.96 |
array/private/copy |
232208 ns |
235708 ns |
0.99 |
array/private/copyto!/cpu_to_gpu |
198875 ns |
199375 ns |
1.00 |
array/private/copyto!/gpu_to_cpu |
199625 ns |
198541 ns |
1.01 |
array/private/copyto!/gpu_to_gpu |
196625 ns |
198333 ns |
0.99 |
array/private/iteration/findall/bool |
1063583 ns |
1067167 ns |
1.00 |
array/private/iteration/findall/int |
1222042 ns |
1224083 ns |
1.00 |
array/private/iteration/findfirst/bool |
1095750 ns |
910833 ns |
1.20 |
array/private/iteration/findfirst/int |
1137708 ns |
1136500 ns |
1.00 |
array/private/iteration/findmin/1d |
1210167 ns |
1221417 ns |
0.99 |
array/private/iteration/findmin/2d |
1053083 ns |
1048917 ns |
1.00 |
array/private/iteration/logical |
1710416 ns |
1708083 ns |
1.00 |
array/private/iteration/scalar |
1228500 ns |
1286125 ns |
0.96 |
array/random/rand/Float32 |
426042 ns |
430833 ns |
0.99 |
array/random/rand/Int64 |
518125 ns |
520666 ns |
1.00 |
array/random/rand!/Float32 |
380625 ns |
384792 ns |
0.99 |
array/random/rand!/Int64 |
422959 ns |
404125 ns |
1.05 |
array/random/randn/Float32 |
400166 ns |
402459 ns |
0.99 |
array/random/randn!/Float32 |
351167 ns |
292542 ns |
1.20 |
array/reductions/mapreduce/Float32/1d |
408291 ns |
424625 ns |
0.96 |
array/reductions/mapreduce/Float32/dims=1 |
338625 ns |
343417 ns |
0.99 |
array/reductions/mapreduce/Float32/dims=1L |
624625 ns |
629792 ns |
0.99 |
array/reductions/mapreduce/Float32/dims=2 |
348916 ns |
351250 ns |
0.99 |
array/reductions/mapreduce/Float32/dims=2L |
1073542 ns |
1062333 ns |
1.01 |
array/reductions/mapreduce/Int64/1d |
543584 ns |
608667 ns |
0.89 |
array/reductions/mapreduce/Int64/dims=1 |
629958 ns |
454208 ns |
1.39 |
array/reductions/mapreduce/Int64/dims=1L |
1039583 ns |
1027542 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=2 |
784833 ns |
785375 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2L |
2189542 ns |
2193208 ns |
1.00 |
array/reductions/reduce/Float32/1d |
433375 ns |
441208 ns |
0.98 |
array/reductions/reduce/Float32/dims=1 |
289083 ns |
290917 ns |
0.99 |
array/reductions/reduce/Float32/dims=1L |
629916 ns |
631750 ns |
1.00 |
array/reductions/reduce/Float32/dims=2 |
233084 ns |
229667 ns |
1.01 |
array/reductions/reduce/Float32/dims=2L |
452959 ns |
450875 ns |
1.00 |
array/reductions/reduce/Int64/1d |
604084 ns |
611417 ns |
0.99 |
array/reductions/reduce/Int64/dims=1 |
459875 ns |
636542 ns |
0.72 |
array/reductions/reduce/Int64/dims=1L |
1032167 ns |
1031291 ns |
1.00 |
array/reductions/reduce/Int64/dims=2 |
243875 ns |
238084 ns |
1.02 |
array/reductions/reduce/Int64/dims=2L |
644417 ns |
646084 ns |
1.00 |
array/shared/copy |
136208 ns |
136292 ns |
1.00 |
array/shared/copyto!/cpu_to_gpu |
37916 ns |
38125 ns |
0.99 |
array/shared/copyto!/gpu_to_cpu |
37625 ns |
37625 ns |
1 |
array/shared/copyto!/gpu_to_gpu |
38000 ns |
38208 ns |
0.99 |
array/shared/iteration/findall/bool |
1068000 ns |
1097000 ns |
0.97 |
array/shared/iteration/findall/int |
1215167 ns |
1232125 ns |
0.99 |
array/shared/iteration/findfirst/bool |
648958 ns |
959417 ns |
0.68 |
array/shared/iteration/findfirst/int |
960375 ns |
675750 ns |
1.42 |
array/shared/iteration/findmin/1d |
1081917 ns |
1078750 ns |
1.00 |
array/shared/iteration/findmin/2d |
1053500 ns |
1053000 ns |
1.00 |
array/shared/iteration/logical |
1564625 ns |
1549833 ns |
1.01 |
array/shared/iteration/scalar |
3744.875 ns |
3765.625 ns |
0.99 |
array/sorting/1d |
2089083 ns |
2117166 ns |
0.99 |
array/sorting/2d |
8308584 ns |
8292708 ns |
1.00 |
integration/byval/reference |
1104125 ns |
1103833 ns |
1.00 |
integration/byval/slices=1 |
1113042 ns |
1106541 ns |
1.01 |
integration/byval/slices=2 |
2020708 ns |
2017500 ns |
1.00 |
integration/byval/slices=3 |
6507459 ns |
6579333 ns |
0.99 |
integration/metaldevrt |
381209 ns |
378417 ns |
1.01 |
kernel/indexing |
184458 ns |
210375 ns |
0.88 |
kernel/indexing_checked |
351875 ns |
372375 ns |
0.94 |
kernel/launch |
1858.3 ns |
1895.8 ns |
0.98 |
kernel/rand |
313208 ns |
373333 ns |
0.84 |
latency/import |
1797781916 ns |
1791487000 ns |
1.00 |
latency/precompile |
26253956291 ns |
26158613125 ns |
1.00 |
latency/ttfp |
2300732333 ns |
2293400750 ns |
1.00 |
metal/synchronization/context |
561.1559139784946 ns |
556 ns |
1.01 |
metal/synchronization/stream |
354.07042253521126 ns |
345.9859813084112 ns |
1.02 |
This comment was automatically generated by workflow using github-action-benchmark.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #965 +/- ##
==========================================
+ Coverage 86.69% 86.71% +0.01%
==========================================
Files 77 77
Lines 5488 5488
==========================================
+ Hits 4758 4759 +1
+ Misses 730 729 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps the compat bounds to LLVMDowngrader_jll 0.11 (JuliaPackaging/Yggdrasil#14758, built against LLVM 23) and GPUCompiler 2.8 (JuliaGPU/GPUCompiler.jl#931). The downgrader C API is unchanged, so this is compat-only.