Skip to content

Upgrade to LLVM.jl 10 - #3323

Merged
maleadt merged 8 commits into
mainfrom
tb/llvm10
Oct 3, 2026
Merged

maleadt merged 8 commits into
mainfrom
tb/llvm10

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Ports CUDA.jl to LLVM.jl 10. Supersedes #3305, whose commit comes first here.

Main changes:

  • Mechanical port: compat bounds, vocabulary imports, properties instead of accessors, and explicit insertion points. Named metadata is created explicitly, which makes the empty pass manager in link_libraries! unnecessary.
  • No more raw C API: LLVM.jl's scoped enums and lookups instead, including for the atomicrmw table.
  • @llvmgenerated for the remaining IR string templates (cluster barriers, grid barrier, RNG state, volumerhs benchmark). The indexing intrinsics are declared with Intrinsic, pointers are converted with pointercast!, the grid barrier uses volatile_load, and atomics pass their syncscope by name. The generated IR is unchanged.

Fixes and behavior changes:

  • compile() inspected the compiled IR after the JuliaContext that owns it was gone. It now inspects and disposes of the IR inside the block.
  • The RNG prologue overwrote the debug location of the kernel's first instruction.
  • The assertion generator's string globals now have fixed names instead of names from a session-wide counter.
  • _pointerref_ldg from Generate IR with LLVM.jl's @llvmgenerated and generate_llvmcall #3305 is now defined on every LLVM version, which fixes the device-function test on older LLVM.
  • The development instructions point at LLVM.jl's main branch.

Tested on an RTX 5080: the full suite on Julia 1.12 and the core tests on 1.10, 1.11 and 1.13 match main. There is one timing-dependent unsafe_wrap failure in core/array.jl, which main also has. LLVM.jl's memcheck mode on 1.12 reports nothing.

Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).


Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.

The indexing, shared memory, malloc, atomics, assertion, printf, ldg and
string pointer intrinsics, and the constant tables for randn, all built
their IR with create_function/call_function around a lot of context and
builder boilerplate. LLVM.jl 9.14's @llvmgenerated and generate_llvmcall
derive the LLVM signature from the Julia one and verify the generated
IR, so use those instead. Julia code that ran before the IR, like
checking the atomic scope or adjusting the index for ldg, moves to a
wrapper that calls the @llvmgenerated function.

The @llvmgenerated helper for ldg is defined on every LLVM version, not
only on LLVM 20 and later where it is used: the device function test in
core/device/method_table collects every @device_function definition
regardless of @static conditions, and failed on older LLVM because the
helper didn't exist there.
Bump the LLVM.jl compat of CUDACore, CUDATools and cuBLAS to 10, and
import the IR and Build vocabularies in CUDACore now that `using LLVM`
only exports @dispose. Accessors and setters become properties, the
builder and new block are positioned with insertion points, named
metadata is created explicitly with get!, and the empty module pass
manager that was only run to attach that metadata is removed.
Refer to linkages, atomic orderings, atomicrmw operations and the code
generation file type through LLVM.jl's scoped enumerations instead of
LLVM.API, and look up atomicrmw operations by their IR name rather than
through a hand-written table. Look up the target globals with get, and
get or declare deferred_codegen with get!.
The cluster barriers were IR
strings that declared the intrinsic with the attributes older LLVM
can't infer, and the volatile load in the grid barrier was an IR string
with a branch for typed pointers. Both are built with @llvmgenerated,
declaring the intrinsic and its attributes with LLVM.jl, which handles
the pointer representation on every LLVM version. The same goes for the
weak shared-memory globals that hold the RNG state, which keep their
names and linkage, and for the arithmetic with contract and nsz flags in
the volumerhs benchmark, which interpolated the operation and type into
IR.

The constant tables of randn were generated by a closure that captured
the table and its element type, which generate_llvmcall compiles again
for every table. Use a named generator that takes the table's name and
element type as statically-known arguments instead, so that it is
compiled once.

The assertion generator named its strings with a global counter, which
made the generated IR depend on session state. The strings are private
globals, whose names LLVM keeps unique, so give them fixed names.
Declare the indexing intrinsics with Intrinsic, which gives them LLVM's
declaration and attributes, looking them up with tryparse(Intrinsic,
name) and falling back to a declaration by name for the cluster
registers that LLVM only knows since version 17. Convert between
pointers and Julia's representation of them with pointercast! rather
than ptrtoint!, and derive the type deferred_codegen returns from
Ptr{Cvoid} instead of spelling out its lowering per version.

Poll the grid barrier with LLVM.Interop's volatile_load instead of a
generator of our own, pass the atomics' synchronization scope by name so
that it resolves in the builder's context, pass the alignment of the
ldg load to load! and refer to the metadata kinds that are part of
LLVM.IR without qualification.
The initialization code that finish_module! prepends to kernels using
the RNG called `debuglocation!(builder, inst)`, meaning to take the
location of the kernel's first instruction, but that copied the
builder's location to the instruction instead: with debug info the
first instruction got a line-zero location, and without it lost its
location. Give the generated instructions the line-zero location, and
leave the kernel's own instructions alone.
compile() returned the IR from the JuliaContext block and then looked
at its functions and at the entry point's name after the block had
returned. That relied on the context leaking (GPUCompiler.jl#970): the
IR is only meant to be used while its context is, and memcheck reports
using it afterwards. Inspect the IR inside the block, and dispose of it
there now that nothing else needs it, so that it doesn't leak the module
and keeps working once that leak is fixed.
@maleadt
maleadt merged commit 3301021 into main Oct 3, 2026
1 check was pending
@maleadt
maleadt deleted the tb/llvm10 branch October 3, 2026 13:01
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 92363d5 Previous: a3202f2 Ratio
array/accumulate/Float32/1d 99648 ns 99028 ns 1.01
array/accumulate/Float32/dims=1 73924 ns 72921 ns 1.01
array/accumulate/Float32/dims=1L 1590742 ns 1588681 ns 1.00
array/accumulate/Float32/dims=2 139360 ns 138806 ns 1.00
array/accumulate/Float32/dims=2L 655811 ns 655618 ns 1.00
array/accumulate/Int64/1d 117734 ns 118327 ns 0.99
array/accumulate/Int64/dims=1 77448 ns 76942 ns 1.01
array/accumulate/Int64/dims=1L 1700146 ns 1697861 ns 1.00
array/accumulate/Int64/dims=2 152265 ns 151829 ns 1.00
array/accumulate/Int64/dims=2L 987515 ns 986767 ns 1.00
array/broadcast 16081 ns 15832 ns 1.02
array/broadcast launch 7253.75 ns 6879.4 ns 1.05
array/construct 870.9814814814815 ns 866.433962264151 ns 1.01
array/copy 16811 ns 16688 ns 1.01
array/copyto!/cpu_to_gpu 210263 ns 208406 ns 1.01
array/copyto!/gpu_to_cpu 241657 ns 241247 ns 1.00
array/copyto!/gpu_to_gpu 9328.333333333334 ns 8802.333333333334 ns 1.06
array/iteration/findall/bool 130764 ns 130199 ns 1.00
array/iteration/findall/int 142387 ns 138662 ns 1.03
array/iteration/findfirst/bool 68474 ns 67562 ns 1.01
array/iteration/findfirst/int 70452 ns 69121 ns 1.02
array/iteration/findmin/1d 63135 ns 59127 ns 1.07
array/iteration/findmin/2d 97431 ns 96956 ns 1.00
array/iteration/logical 185792 ns 180975 ns 1.03
array/iteration/scalar 59908 ns 58859 ns 1.02
array/permutedims/2d 46149 ns 46242 ns 1.00
array/permutedims/3d 47937 ns 46855 ns 1.02
array/permutedims/4d 48780 ns 48368 ns 1.01
array/random/rand/Float32 11127 ns 11809 ns 0.94
array/random/rand/Int64 18595 ns 18752 ns 0.99
array/random/rand!/Float32 7684.25 ns 7829.5 ns 0.98
array/random/rand!/Int64 16959 ns 15615 ns 1.09
array/random/randn/Float32 32537 ns 32695 ns 1.00
array/random/randn!/Float32 23377 ns 23566 ns 0.99
array/reductions/mapreduce/Float32/1d 33358 ns 31821 ns 1.05
array/reductions/mapreduce/Float32/dims=1 37784 ns 37330 ns 1.01
array/reductions/mapreduce/Float32/dims=1L 51487 ns 50978 ns 1.01
array/reductions/mapreduce/Float32/dims=2 55805 ns 54829 ns 1.02
array/reductions/mapreduce/Float32/dims=2L 67958 ns 67220 ns 1.01
array/reductions/mapreduce/Int64/1d 39339 ns 39331 ns 1.00
array/reductions/mapreduce/Int64/dims=1 40935 ns 40315 ns 1.02
array/reductions/mapreduce/Int64/dims=1L 89141 ns 88477 ns 1.01
array/reductions/mapreduce/Int64/dims=2 58259 ns 57574 ns 1.01
array/reductions/mapreduce/Int64/dims=2L 84854 ns 83461 ns 1.02
array/reductions/reduce/Float32/1d 33547 ns 32196 ns 1.04
array/reductions/reduce/Float32/dims=1 37906 ns 37272 ns 1.02
array/reductions/reduce/Float32/dims=1L 51247 ns 50697 ns 1.01
array/reductions/reduce/Float32/dims=2 55739 ns 55042 ns 1.01
array/reductions/reduce/Float32/dims=2L 68153 ns 67551 ns 1.01
array/reductions/reduce/Int64/1d 39590 ns 39026 ns 1.01
array/reductions/reduce/Int64/dims=1 40738 ns 40114 ns 1.02
array/reductions/reduce/Int64/dims=1L 88973 ns 88494 ns 1.01
array/reductions/reduce/Int64/dims=2 58342 ns 57350 ns 1.02
array/reductions/reduce/Int64/dims=2L 84261 ns 83169 ns 1.01
array/reverse/1d 17030 ns 17038 ns 1.00
array/reverse/1dL 70385 ns 69729 ns 1.01
array/reverse/1dL_inplace 67892 ns 67315 ns 1.01
array/reverse/1d_inplace 8899 ns 8467.666666666666 ns 1.05
array/reverse/2d 20424 ns 20236 ns 1.01
array/reverse/2dL 73947 ns 73577 ns 1.01
array/reverse/2dL_inplace 67493 ns 67185 ns 1.00
array/reverse/2d_inplace 10127 ns 9814 ns 1.03
array/sorting/1d 2641364 ns 2646705 ns 1.00
array/sorting/2d 1019204 ns 1017870 ns 1.00
array/sorting/by 3175927 ns 3174206 ns 1.00
cuda/synchronization/context/auto 6804.2 ns 6849.6 ns 0.99
cuda/synchronization/context/blocking 795.7368421052631 ns 815.9772727272727 ns 0.98
cuda/synchronization/context/nonblocking 6993.5 ns 6897.4 ns 1.01
cuda/synchronization/stream/auto 697.2237762237762 ns 700.5918367346939 ns 1.00
cuda/synchronization/stream/blocking 870.0181818181818 ns 892.3673469387755 ns 0.97
cuda/synchronization/stream/nonblocking 7376.5 ns 7433 ns 0.99
integration/byval/reference 148375 ns 148108 ns 1.00
integration/byval/slices=1 149537 ns 149157 ns 1.00
integration/byval/slices=2 292636 ns 291918 ns 1.00
integration/byval/slices=3 435819 ns 434943 ns 1.00
integration/cudadevrt 105383 ns 105211 ns 1.00
integration/volumerhs 9147261 ns 9139091 ns 1.00
kernel/indexing 13945 ns 13078 ns 1.07
kernel/indexing_checked 14121 ns 13558 ns 1.04
kernel/launch 2347.777777777778 ns 2086.222222222222 ns 1.13
kernel/occupancy 935.6153846153846 ns 684.6734693877551 ns 1.37
kernel/rand 14772 ns 16385 ns 0.90
latency/import 4325261524 ns 4265954564 ns 1.01
latency/precompile 5065611849 ns 5038240985 ns 1.01
latency/ttfp 4827529727 ns 4715207932 ns 1.02

This comment was automatically generated by workflow using github-action-benchmark.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant