Skip to content

Upgrade to LLVM.jl 10 - #997

Merged
maleadt merged 4 commits into
mainfrom
tb/llvm10
Oct 4, 2026
Merged

maleadt merged 4 commits into
mainfrom
tb/llvm10

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Ports Metal.jl to LLVM.jl 10.

Main changes:

  • Mechanical port: compat, vocabulary imports, properties instead of accessors, and explicit insertion points. The threadgroup memory, RNG state and @mtlprintf generators use @llvmgenerated.
  • No more raw C API: scoped enums, get/get! on the module views, llvm.memcpy recognized as an intrinsic rather than by its name prefix, and strip_pointer_casts instead of hand-written pointer walks. @mtlprintf declares the varargs and lifetime intrinsics through Intrinsic, which handles their changes in LLVM 19 and 22. The volumerhs benchmark uses @llvmgenerated.
  • IR ownership: the IR that GPUCompiler returns is disposed of inside the JuliaContext block, also when lowering to AIR fails.

Fixes:

  • finish_module! overwrote the debug location of the kernel's first instruction with a line-0 location, and the AIR downgrades copied a call's location onto itself.
  • The simdgroup-matrix downgrade passed non-constant vectors to LLVMGetAggregateElement, which asserts on LLVM builds with assertions.

Tested on an M1 on Julia 1.10, 1.12 and 1.13, including GPUArrays' tests. Results match main: on 1.10, the two fence checks in synchronization.jl fail as they do on main, and the gtk example needs a window server.

Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).


Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.

Bump the LLVM.jl compat to 10, import the IR and Build vocabularies next
to LLVM.Interop (`using LLVM` no longer exports the API), and use
properties instead of the removed accessor functions: module functions
and globals, names, linkage, initializers, alignment, sections,
metadata, function attributes, operands, arguments, users, function
types, and so on. The bin scripts and the scripts test import LLVM.IR
too.

`create_function`/`call_function` were removed: the threadgroup memory,
RNG state and `_mtlprintf` generators become `@llvmgenerated` functions.
The latter builds the `metal_os_log` helper with a scoped `position!`,
takes the types of its varargs from their LLVM values, and computes
their sizes with `LLVM.storage_size(dl, T)` instead of `sizeof(dl, T)`
(an `Int` rather than a `Float64`).

Insertion points are explicit: the RNG initialization block is created
before the entry block with `LLVM.before`, and the AIR downgrades
position their builder `LLVM.before` the call they rewrite.

The `debuglocation!(builder, inst)` calls copied the builder's debug
location to the instruction, contrary to what its documentation said:
- `finish_module!` used it on the entry block's first instruction after
  giving the builder a line-0 location, which overwrote (or, without a
  subprogram, cleared) that instruction's location. The intent was to
  give the RNG initialization prologue a location, so only set the
  builder's location now, leaving the original instructions alone.
- `finish_ir!` used it right after positioning the builder before the
  call being replaced, which already gives the builder the call's
  location, so it copied the call's location onto itself. Drop it.
…ngrade

`static_vector_lane` passed any value that isn't an `insertelement` to
`LLVMGetAggregateElement`, e.g. a strides vector that is loaded or comes
from a phi or an argument. That function casts its argument to a
`Constant` with LLVM's checked `cast`, which asserts on builds of LLVM
with assertions, and only happens to return null (not a known lane) on
release builds. Return `nothing` for non-constants without calling it.
Name linkages and the code generation file type with LLVM.jl's scoped
enumerations, key the set of tensor-op descriptor allocas on the values
themselves instead of their handles, and read the lanes of constant
vectors with their `elements`.

Look up or declare globals and functions with `get`/`get!` on the
module's views, copy the attributes of redeclared AIR intrinsics with
`append!`, and recognize `llvm.memcpy` and `llvm.memcpy.inline` calls by
their intrinsics instead of a name prefix. The hand-written walks from a
pointer to the alloca or global it is based on become
`strip_pointer_casts`, which strips the same casts and all-zero GEPs,
but also when they are constant expressions (the old walks only looked
through instructions).

Declare the intrinsics that `_mtlprintf` uses through LLVM.jl's
`Intrinsic`s instead of by mangled names with typed pointers, which only
worked because `llvmcall` upgrades outdated declarations. This also
handles `llvm.va_start`/`llvm.va_end` being overloaded since LLVM 19,
and the lifetime intrinsics having lost their size operand in LLVM 22.

Also derive the type of the `deferred_codegen` declaration from Julia's
lowering of `Ptr{Cvoid}`, like GPUCompiler declares it, instead of a
version decision tree, set the volatile bit of AIR atomics through the
call's `arguments` instead of indexing its operands (where the callee
comes last), and create attributes from their kind's symbol.

The volumerhs benchmark's arithmetic with fast-math flags interpolated
the operation into IR strings; generate it with `@llvmgenerated`, setting
the flags through the instruction's `fast_math` property (the benchmark
environment now depends on LLVM.jl for that).
GPUCompiler now hands the module that Julia's code generator produced to
the caller of `compile` (on Julia 1.11 and later, it moves it out of the
native code), so the IR that `compile_to_metallib` lowers to AIR, and
that the relocations test only compiles to inspect its relocations,
belongs to us. Dispose of it inside the `JuliaContext` block once it has
been lowered, printed and its entry point's name read, rather than
leaking it until the context is disposed of. `compile_to_metallib` does
so with `@dispose`, so that the module is also disposed of when lowering
it to AIR, or dumping it after that failed, throws.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Metal Benchmarks

Details
Benchmark suite Current: 440e539 Previous: 6eb4996 Ratio
array/accumulate/Float32/1d 388542 ns 393500 ns 0.99
array/accumulate/Float32/dims=1 366875 ns 368750 ns 0.99
array/accumulate/Float32/dims=1L 8847042 ns 8879000 ns 1.00
array/accumulate/Float32/dims=2 431500 ns 436667 ns 0.99
array/accumulate/Float32/dims=2L 2640958 ns 2760209 ns 0.96
array/accumulate/Int64/1d 842625 ns 849667 ns 0.99
array/accumulate/Int64/dims=1 908917 ns 900750 ns 1.01
array/accumulate/Int64/dims=1L 9553709 ns 9583625 ns 1.00
array/accumulate/Int64/dims=2 1210084 ns 1215250 ns 1.00
array/accumulate/Int64/dims=2L 6557542 ns 6538208 ns 1.00
array/broadcast 227958 ns 237708 ns 0.96
array/construct 2292 ns 2292 ns 1
array/permutedims/2d 458666 ns 444000 ns 1.03
array/permutedims/3d 1030542 ns 1029208 ns 1.00
array/permutedims/4d 1115167 ns 1119959 ns 1.00
array/private/copy 229459 ns 228417 ns 1.00
array/private/copyto!/cpu_to_gpu 213709 ns 213417 ns 1.00
array/private/copyto!/gpu_to_cpu 216458 ns 210208 ns 1.03
array/private/copyto!/gpu_to_gpu 220542 ns 219250 ns 1.01
array/private/iteration/findall/bool 1061667 ns 1056458 ns 1.00
array/private/iteration/findall/int 1224875 ns 1217917 ns 1.01
array/private/iteration/findfirst/bool 1155667 ns 1154084 ns 1.00
array/private/iteration/findfirst/int 1161042 ns 1158833 ns 1.00
array/private/iteration/findmin/1d 1193458 ns 1191875 ns 1.00
array/private/iteration/findmin/2d 1028542 ns 1022417 ns 1.01
array/private/iteration/logical 1667166 ns 1663292 ns 1.00
array/private/iteration/scalar 1406500 ns 1380333 ns 1.02
array/random/rand/Float32 419625 ns 418875 ns 1.00
array/random/rand/Int64 493167 ns 494333 ns 1.00
array/random/rand!/Float32 379292 ns 403916 ns 0.94
array/random/rand!/Int64 427375 ns 391417 ns 1.09
array/random/randn/Float32 393833 ns 384167 ns 1.03
array/random/randn!/Float32 377375 ns 368417 ns 1.02
array/reductions/mapreduce/Float32/1d 450709 ns 438709 ns 1.03
array/reductions/mapreduce/Float32/dims=1 356417 ns 290625 ns 1.23
array/reductions/mapreduce/Float32/dims=1L 617459 ns 618333 ns 1.00
array/reductions/mapreduce/Float32/dims=2 358583 ns 349250 ns 1.03
array/reductions/mapreduce/Float32/dims=2L 899166 ns 1094500 ns 0.82
array/reductions/mapreduce/Int64/1d 621042 ns 636292 ns 0.98
array/reductions/mapreduce/Int64/dims=1 637125 ns 631500 ns 1.01
array/reductions/mapreduce/Int64/dims=1L 1007583 ns 1010458 ns 1.00
array/reductions/mapreduce/Int64/dims=2 791875 ns 785625 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 2191167 ns 2189584 ns 1.00
array/reductions/reduce/Float32/1d 449625 ns 449375 ns 1.00
array/reductions/reduce/Float32/dims=1 360708 ns 355666 ns 1.01
array/reductions/reduce/Float32/dims=1L 616667 ns 607917 ns 1.01
array/reductions/reduce/Float32/dims=2 246291 ns 249042 ns 0.99
array/reductions/reduce/Float32/dims=2L 463250 ns 473250 ns 0.98
array/reductions/reduce/Int64/1d 638792 ns 629875 ns 1.01
array/reductions/reduce/Int64/dims=1 637084 ns 637833 ns 1.00
array/reductions/reduce/Int64/dims=1L 1018959 ns 1005875 ns 1.01
array/reductions/reduce/Int64/dims=2 256791 ns 268709 ns 0.96
array/reductions/reduce/Int64/dims=2L 672542 ns 666541 ns 1.01
array/shared/copy 129917 ns 131292 ns 0.99
array/shared/copyto!/cpu_to_gpu 38334 ns 38042 ns 1.01
array/shared/copyto!/gpu_to_cpu 37959 ns 37292 ns 1.02
array/shared/copyto!/gpu_to_gpu 38250 ns 37625 ns 1.02
array/shared/iteration/findall/bool 1062750 ns 1045500 ns 1.02
array/shared/iteration/findall/int 1220291 ns 1221500 ns 1.00
array/shared/iteration/findfirst/bool 981459 ns 977500 ns 1.00
array/shared/iteration/findfirst/int 991709 ns 986500 ns 1.01
array/shared/iteration/findmin/1d 1054333 ns 1066708 ns 0.99
array/shared/iteration/findmin/2d 1033375 ns 1030959 ns 1.00
array/shared/iteration/logical 1525292 ns 1530416 ns 1.00
array/shared/iteration/scalar 4666.714285714285 ns 4875 ns 0.96
array/sorting/1d 2108250 ns 2093791 ns 1.01
array/sorting/2d 8431166 ns 8467208 ns 1.00
integration/byval/reference 1105250 ns 1123250 ns 0.98
integration/byval/slices=1 1127125 ns 1116541 ns 1.01
integration/byval/slices=2 2028750 ns 2019584 ns 1.00
integration/byval/slices=3 6846292 ns 6809084 ns 1.01
integration/metaldevrt 400042 ns 392417 ns 1.02
kernel/indexing 216708 ns 208833 ns 1.04
kernel/indexing_checked 394500 ns 393834 ns 1.00
kernel/launch 2074 ns 2111.1111111111113 ns 0.98
kernel/rand 402084 ns 403041 ns 1.00
latency/import 1860023458 ns 1821483458 ns 1.02
latency/precompile 33629446292 ns 32771377958 ns 1.03
latency/ttfp 2287592542 ns 2276943625 ns 1.00
metal/synchronization/context 528.7315789473685 ns 518.7643979057592 ns 1.02
metal/synchronization/stream 439.8131313131313 ns 459.39086294416245 ns 0.96

This comment was automatically generated by workflow using github-action-benchmark.

@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.52941% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 87.23%. Comparing base (6eb4996) to head (440e539).

Files with missing lines Patch % Lines
src/compiler/compilation.jl 97.22% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #997      +/-   ##
==========================================
+ Coverage   87.09%   87.23%   +0.14%     
==========================================
  Files          92       92              
  Lines        6747     6666      -81     
==========================================
- Hits         5876     5815      -61     
+ Misses        871      851      -20     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt merged commit 83e8268 into main Oct 4, 2026
19 checks passed
@maleadt
maleadt deleted the tb/llvm10 branch October 4, 2026 05:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant