Conversation
Bump the LLVM.jl compat to 10, import the IR and Build vocabularies next to LLVM.Interop (`using LLVM` no longer exports the API), and use properties instead of the removed accessor functions: module functions and globals, names, linkage, initializers, alignment, sections, metadata, function attributes, operands, arguments, users, function types, and so on. The bin scripts and the scripts test import LLVM.IR too. `create_function`/`call_function` were removed: the threadgroup memory, RNG state and `_mtlprintf` generators become `@llvmgenerated` functions. The latter builds the `metal_os_log` helper with a scoped `position!`, takes the types of its varargs from their LLVM values, and computes their sizes with `LLVM.storage_size(dl, T)` instead of `sizeof(dl, T)` (an `Int` rather than a `Float64`). Insertion points are explicit: the RNG initialization block is created before the entry block with `LLVM.before`, and the AIR downgrades position their builder `LLVM.before` the call they rewrite. The `debuglocation!(builder, inst)` calls copied the builder's debug location to the instruction, contrary to what its documentation said: - `finish_module!` used it on the entry block's first instruction after giving the builder a line-0 location, which overwrote (or, without a subprogram, cleared) that instruction's location. The intent was to give the RNG initialization prologue a location, so only set the builder's location now, leaving the original instructions alone. - `finish_ir!` used it right after positioning the builder before the call being replaced, which already gives the builder the call's location, so it copied the call's location onto itself. Drop it.
…ngrade `static_vector_lane` passed any value that isn't an `insertelement` to `LLVMGetAggregateElement`, e.g. a strides vector that is loaded or comes from a phi or an argument. That function casts its argument to a `Constant` with LLVM's checked `cast`, which asserts on builds of LLVM with assertions, and only happens to return null (not a known lane) on release builds. Return `nothing` for non-constants without calling it.
Name linkages and the code generation file type with LLVM.jl's scoped
enumerations, key the set of tensor-op descriptor allocas on the values
themselves instead of their handles, and read the lanes of constant
vectors with their `elements`.
Look up or declare globals and functions with `get`/`get!` on the
module's views, copy the attributes of redeclared AIR intrinsics with
`append!`, and recognize `llvm.memcpy` and `llvm.memcpy.inline` calls by
their intrinsics instead of a name prefix. The hand-written walks from a
pointer to the alloca or global it is based on become
`strip_pointer_casts`, which strips the same casts and all-zero GEPs,
but also when they are constant expressions (the old walks only looked
through instructions).
Declare the intrinsics that `_mtlprintf` uses through LLVM.jl's
`Intrinsic`s instead of by mangled names with typed pointers, which only
worked because `llvmcall` upgrades outdated declarations. This also
handles `llvm.va_start`/`llvm.va_end` being overloaded since LLVM 19,
and the lifetime intrinsics having lost their size operand in LLVM 22.
Also derive the type of the `deferred_codegen` declaration from Julia's
lowering of `Ptr{Cvoid}`, like GPUCompiler declares it, instead of a
version decision tree, set the volatile bit of AIR atomics through the
call's `arguments` instead of indexing its operands (where the callee
comes last), and create attributes from their kind's symbol.
The volumerhs benchmark's arithmetic with fast-math flags interpolated
the operation into IR strings; generate it with `@llvmgenerated`, setting
the flags through the instruction's `fast_math` property (the benchmark
environment now depends on LLVM.jl for that).
GPUCompiler now hands the module that Julia's code generator produced to the caller of `compile` (on Julia 1.11 and later, it moves it out of the native code), so the IR that `compile_to_metallib` lowers to AIR, and that the relocations test only compiles to inspect its relocations, belongs to us. Dispose of it inside the `JuliaContext` block once it has been lowered, printed and its entry point's name read, rather than leaking it until the context is disposed of. `compile_to_metallib` does so with `@dispose`, so that the module is also disposed of when lowering it to AIR, or dumping it after that failed, throws.
Contributor
There was a problem hiding this comment.
Metal Benchmarks
Details
| Benchmark suite | Current: 440e539 | Previous: 6eb4996 | Ratio |
|---|---|---|---|
array/accumulate/Float32/1d |
388542 ns |
393500 ns |
0.99 |
array/accumulate/Float32/dims=1 |
366875 ns |
368750 ns |
0.99 |
array/accumulate/Float32/dims=1L |
8847042 ns |
8879000 ns |
1.00 |
array/accumulate/Float32/dims=2 |
431500 ns |
436667 ns |
0.99 |
array/accumulate/Float32/dims=2L |
2640958 ns |
2760209 ns |
0.96 |
array/accumulate/Int64/1d |
842625 ns |
849667 ns |
0.99 |
array/accumulate/Int64/dims=1 |
908917 ns |
900750 ns |
1.01 |
array/accumulate/Int64/dims=1L |
9553709 ns |
9583625 ns |
1.00 |
array/accumulate/Int64/dims=2 |
1210084 ns |
1215250 ns |
1.00 |
array/accumulate/Int64/dims=2L |
6557542 ns |
6538208 ns |
1.00 |
array/broadcast |
227958 ns |
237708 ns |
0.96 |
array/construct |
2292 ns |
2292 ns |
1 |
array/permutedims/2d |
458666 ns |
444000 ns |
1.03 |
array/permutedims/3d |
1030542 ns |
1029208 ns |
1.00 |
array/permutedims/4d |
1115167 ns |
1119959 ns |
1.00 |
array/private/copy |
229459 ns |
228417 ns |
1.00 |
array/private/copyto!/cpu_to_gpu |
213709 ns |
213417 ns |
1.00 |
array/private/copyto!/gpu_to_cpu |
216458 ns |
210208 ns |
1.03 |
array/private/copyto!/gpu_to_gpu |
220542 ns |
219250 ns |
1.01 |
array/private/iteration/findall/bool |
1061667 ns |
1056458 ns |
1.00 |
array/private/iteration/findall/int |
1224875 ns |
1217917 ns |
1.01 |
array/private/iteration/findfirst/bool |
1155667 ns |
1154084 ns |
1.00 |
array/private/iteration/findfirst/int |
1161042 ns |
1158833 ns |
1.00 |
array/private/iteration/findmin/1d |
1193458 ns |
1191875 ns |
1.00 |
array/private/iteration/findmin/2d |
1028542 ns |
1022417 ns |
1.01 |
array/private/iteration/logical |
1667166 ns |
1663292 ns |
1.00 |
array/private/iteration/scalar |
1406500 ns |
1380333 ns |
1.02 |
array/random/rand/Float32 |
419625 ns |
418875 ns |
1.00 |
array/random/rand/Int64 |
493167 ns |
494333 ns |
1.00 |
array/random/rand!/Float32 |
379292 ns |
403916 ns |
0.94 |
array/random/rand!/Int64 |
427375 ns |
391417 ns |
1.09 |
array/random/randn/Float32 |
393833 ns |
384167 ns |
1.03 |
array/random/randn!/Float32 |
377375 ns |
368417 ns |
1.02 |
array/reductions/mapreduce/Float32/1d |
450709 ns |
438709 ns |
1.03 |
array/reductions/mapreduce/Float32/dims=1 |
356417 ns |
290625 ns |
1.23 |
array/reductions/mapreduce/Float32/dims=1L |
617459 ns |
618333 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2 |
358583 ns |
349250 ns |
1.03 |
array/reductions/mapreduce/Float32/dims=2L |
899166 ns |
1094500 ns |
0.82 |
array/reductions/mapreduce/Int64/1d |
621042 ns |
636292 ns |
0.98 |
array/reductions/mapreduce/Int64/dims=1 |
637125 ns |
631500 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=1L |
1007583 ns |
1010458 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2 |
791875 ns |
785625 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=2L |
2191167 ns |
2189584 ns |
1.00 |
array/reductions/reduce/Float32/1d |
449625 ns |
449375 ns |
1.00 |
array/reductions/reduce/Float32/dims=1 |
360708 ns |
355666 ns |
1.01 |
array/reductions/reduce/Float32/dims=1L |
616667 ns |
607917 ns |
1.01 |
array/reductions/reduce/Float32/dims=2 |
246291 ns |
249042 ns |
0.99 |
array/reductions/reduce/Float32/dims=2L |
463250 ns |
473250 ns |
0.98 |
array/reductions/reduce/Int64/1d |
638792 ns |
629875 ns |
1.01 |
array/reductions/reduce/Int64/dims=1 |
637084 ns |
637833 ns |
1.00 |
array/reductions/reduce/Int64/dims=1L |
1018959 ns |
1005875 ns |
1.01 |
array/reductions/reduce/Int64/dims=2 |
256791 ns |
268709 ns |
0.96 |
array/reductions/reduce/Int64/dims=2L |
672542 ns |
666541 ns |
1.01 |
array/shared/copy |
129917 ns |
131292 ns |
0.99 |
array/shared/copyto!/cpu_to_gpu |
38334 ns |
38042 ns |
1.01 |
array/shared/copyto!/gpu_to_cpu |
37959 ns |
37292 ns |
1.02 |
array/shared/copyto!/gpu_to_gpu |
38250 ns |
37625 ns |
1.02 |
array/shared/iteration/findall/bool |
1062750 ns |
1045500 ns |
1.02 |
array/shared/iteration/findall/int |
1220291 ns |
1221500 ns |
1.00 |
array/shared/iteration/findfirst/bool |
981459 ns |
977500 ns |
1.00 |
array/shared/iteration/findfirst/int |
991709 ns |
986500 ns |
1.01 |
array/shared/iteration/findmin/1d |
1054333 ns |
1066708 ns |
0.99 |
array/shared/iteration/findmin/2d |
1033375 ns |
1030959 ns |
1.00 |
array/shared/iteration/logical |
1525292 ns |
1530416 ns |
1.00 |
array/shared/iteration/scalar |
4666.714285714285 ns |
4875 ns |
0.96 |
array/sorting/1d |
2108250 ns |
2093791 ns |
1.01 |
array/sorting/2d |
8431166 ns |
8467208 ns |
1.00 |
integration/byval/reference |
1105250 ns |
1123250 ns |
0.98 |
integration/byval/slices=1 |
1127125 ns |
1116541 ns |
1.01 |
integration/byval/slices=2 |
2028750 ns |
2019584 ns |
1.00 |
integration/byval/slices=3 |
6846292 ns |
6809084 ns |
1.01 |
integration/metaldevrt |
400042 ns |
392417 ns |
1.02 |
kernel/indexing |
216708 ns |
208833 ns |
1.04 |
kernel/indexing_checked |
394500 ns |
393834 ns |
1.00 |
kernel/launch |
2074 ns |
2111.1111111111113 ns |
0.98 |
kernel/rand |
402084 ns |
403041 ns |
1.00 |
latency/import |
1860023458 ns |
1821483458 ns |
1.02 |
latency/precompile |
33629446292 ns |
32771377958 ns |
1.03 |
latency/ttfp |
2287592542 ns |
2276943625 ns |
1.00 |
metal/synchronization/context |
528.7315789473685 ns |
518.7643979057592 ns |
1.02 |
metal/synchronization/stream |
439.8131313131313 ns |
459.39086294416245 ns |
0.96 |
This comment was automatically generated by workflow using github-action-benchmark.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #997 +/- ##
==========================================
+ Coverage 87.09% 87.23% +0.14%
==========================================
Files 92 92
Lines 6747 6666 -81
==========================================
- Hits 5876 5815 -61
+ Misses 871 851 -20 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports Metal.jl to LLVM.jl 10.
Main changes:
@mtlprintfgenerators use@llvmgenerated.get/get!on the module views,llvm.memcpyrecognized as an intrinsic rather than by its name prefix, andstrip_pointer_castsinstead of hand-written pointer walks.@mtlprintfdeclares the varargs and lifetime intrinsics throughIntrinsic, which handles their changes in LLVM 19 and 22. The volumerhs benchmark uses@llvmgenerated.JuliaContextblock, also when lowering to AIR fails.Fixes:
finish_module!overwrote the debug location of the kernel's first instruction with a line-0 location, and the AIR downgrades copied a call's location onto itself.LLVMGetAggregateElement, which asserts on LLVM builds with assertions.Tested on an M1 on Julia 1.10, 1.12 and 1.13, including GPUArrays' tests. Results match main: on 1.10, the two fence checks in
synchronization.jlfail as they do on main, and thegtkexample needs a window server.Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).
Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.