Skip to content

Upgrade to LLVM.jl 10 - #1132

Merged
maleadt merged 5 commits into
mainfrom
tb/llvm10
Oct 6, 2026
Merged

maleadt merged 5 commits into
mainfrom
tb/llvm10

Conversation

@maleadt

@maleadt maleadt commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Ports AMDGPU.jl to LLVM.jl 10.

Main changes:

  • Mechanical port: vocabulary imports, properties instead of accessors, and explicit insertion points. Every hand-written IR generator (indices, LDS and scratch allocation, RNG and Ziggurat tables, device strings, memcpy!/memset!, the exception CAS) is now an @llvmgenerated function. The unused get_global_pointer and string_length were removed instead of ported.
  • No more raw C API: scoped enums, attribute sets keyed by kind, and Intrinsic declarations. The fence overlays use fence! instead of IR strings, and device strings go directly into the global address space.

Fixes, each in its own commit:

  • Zero-initialized LDS arrays were cleared with wrong sizes, often writing past the allocation. For example, @ROCStaticLocalArray(Int32, 256) cleared 4096 bytes instead of 1024. The size now comes from the data layout.
  • hipcompile inspected the compiled IR after the JuliaContext that owns it was gone. It now inspects and disposes of it inside the block.
  • finish_ir! leaked a TargetMachine for every kernel.

The LLVM IR and GCN of 12 kernels covering every ported generator are identical to main's on gfx1030 and gfx90a, apart from the intended fixes. The test suite only ran on a gfx1036 iGPU, which is unsupported and flaky. There, the port does no worse than main (15 failures and 1 error on Julia 1.12, against 187 failures and 111 errors for main). A CI run on a supported GPU is still needed.

Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).


Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.

LLVM.jl 10 reorganizes its API: `using LLVM` only exports `@dispose`, the
rest comes from the `LLVM.IR`, `LLVM.Build` and `LLVM.Passes` vocabularies,
and what an object has is now a property (`f.name`, `mod.functions`,
`gv.linkage = ...`) instead of an accessor function.

- Compiler imports `LLVM.IR`, `LLVM.Build` and `LLVM.Passes`; Device imports
  `LLVM.IR` and `LLVM.Build` (its own `memcpy!`, `memset!` and `free!` shadow
  the builder functions of the same name, which it calls qualified).
- The GPUCompiler hooks, the device library linker and the LDS
  zero-initialization use properties, `LLVM.before` insertion points, the
  public floating-point type names and `PassBuilder`.
- `create_function`/`call_function` were removed, so the IR generators
  (indexing, LDS/scratch allocation, random state and Ziggurat tables, device
  strings, `memcpy!`/`memset!`, the exception CAS) are now `@llvmgenerated`
  functions. They generate the same IR. `emit_constant_array` takes the name
  of the `Random` table and its element type and looks the table up in the
  generator, so the `gpu_ki()`-style accessors became plain functions.
- `get_global_pointer` and `string_length` had no callers, so they were
  removed instead of ported.
Replace the remaining LLVM.API enum constants with LLVM.jl's scoped enums
(LLVM.CallConv, LLVM.Linkage, LLVM.AtomicOrdering, LLVM.AtomicRMWBinOp), and
the hand-written lookups with what LLVM.jl offers: attribute sets keyed by
kind instead of scanning them for the noinline/alwaysinline attributes,
get(mod.functions, ...) and the users of a value in fold_wavefrontsize!, and
intrinsic declarations through Intrinsic instead of by name with a
hand-written function type. The generated IR is unchanged.

The UnsafeAtomics.fence overlays for the agent and workgroup scopes
interpolated UnsafeAtomics' ordering and scope objects into an IR string;
generate the fence with fence! instead, which takes the scope by name.
Device strings are emitted directly in the global address space using
globalstring_ptr!'s addrspace keyword, instead of as a generic global that
is cast to the global address space. isghosttype no longer calls into
Julia's code generator, so the kernel launch generator doesn't need a
temporary LLVM context around it.
zeroinit_lds! computed the number of bytes to clear with a hand-written
llvmsize, which got most types wrong:

- every integer type was 16 bytes, because it took the width of an i128
  GenericValue (which it also leaked) instead of the type's, so integer and
  Bool arrays were cleared with a memset 2-16 times their size, writing past
  the end of the global: @ROCStaticLocalArray(Int32, 256) cleared 4096 bytes
  instead of 1024, and the RNG's per-warp state 512 bytes instead of 128 for
  each of its two arrays;
- vector types were sized by their number of elements instead of bytes, so
  arrays of e.g. NTuple{4,VecElement{Float32}} were only partially cleared
  (32 of 128 bytes for 8 elements);
- non-packed structs were assumed to take 8 bytes per field.

Use LLVM.storage_size with the module's data layout instead.
hipcompile looked for late global hostcalls and externally initialized
globals in meta.ir, and read the entry point's name, after the
JuliaContext() block had returned, i.e. after the thread-safe context the
module belongs to had been disposed of. That relied on the context leaking
(GPUCompiler.jl#970): memcheck reports it as a use of the module after its
context's disposal, and it would stop working once that leak is fixed.

Do this inside the block. The caller owns the module GPUCompiler returns,
so also dispose of it there, like GPUCompiler's own callers of compile do;
otherwise the module leaks along with the context. The codegen test that
calls compile directly does the same.
finish_ir! created a TargetMachine for every kernel it compiled to run the
amdgpu-attributor pass with, and never disposed of it.
@maleadt
maleadt merged commit 2a4ee27 into main Oct 6, 2026
2 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant