Repository navigation
Conversation
maleadt
force-pushed
the
tb/llvm10
branch
2 times, most recently
from
October 3, 2026 09:18
6419791 to
83fda2d
Compare
maleadt
added this pull request to stack #1136
October 5, 2026 20:07
LLVM.jl 10 reorganizes its API: `using LLVM` only exports `@dispose`, the rest comes from the `LLVM.IR`, `LLVM.Build` and `LLVM.Passes` vocabularies, and what an object has is now a property (`f.name`, `mod.functions`, `gv.linkage = ...`) instead of an accessor function. - Compiler imports `LLVM.IR`, `LLVM.Build` and `LLVM.Passes`; Device imports `LLVM.IR` and `LLVM.Build` (its own `memcpy!`, `memset!` and `free!` shadow the builder functions of the same name, which it calls qualified). - The GPUCompiler hooks, the device library linker and the LDS zero-initialization use properties, `LLVM.before` insertion points, the public floating-point type names and `PassBuilder`. - `create_function`/`call_function` were removed, so the IR generators (indexing, LDS/scratch allocation, random state and Ziggurat tables, device strings, `memcpy!`/`memset!`, the exception CAS) are now `@llvmgenerated` functions. They generate the same IR. `emit_constant_array` takes the name of the `Random` table and its element type and looks the table up in the generator, so the `gpu_ki()`-style accessors became plain functions. - `get_global_pointer` and `string_length` had no callers, so they were removed instead of ported.
Replace the remaining LLVM.API enum constants with LLVM.jl's scoped enums (LLVM.CallConv, LLVM.Linkage, LLVM.AtomicOrdering, LLVM.AtomicRMWBinOp), and the hand-written lookups with what LLVM.jl offers: attribute sets keyed by kind instead of scanning them for the noinline/alwaysinline attributes, get(mod.functions, ...) and the users of a value in fold_wavefrontsize!, and intrinsic declarations through Intrinsic instead of by name with a hand-written function type. The generated IR is unchanged. The UnsafeAtomics.fence overlays for the agent and workgroup scopes interpolated UnsafeAtomics' ordering and scope objects into an IR string; generate the fence with fence! instead, which takes the scope by name. Device strings are emitted directly in the global address space using globalstring_ptr!'s addrspace keyword, instead of as a generic global that is cast to the global address space. isghosttype no longer calls into Julia's code generator, so the kernel launch generator doesn't need a temporary LLVM context around it.
zeroinit_lds! computed the number of bytes to clear with a hand-written
llvmsize, which got most types wrong:
- every integer type was 16 bytes, because it took the width of an i128
GenericValue (which it also leaked) instead of the type's, so integer and
Bool arrays were cleared with a memset 2-16 times their size, writing past
the end of the global: @ROCStaticLocalArray(Int32, 256) cleared 4096 bytes
instead of 1024, and the RNG's per-warp state 512 bytes instead of 128 for
each of its two arrays;
- vector types were sized by their number of elements instead of bytes, so
arrays of e.g. NTuple{4,VecElement{Float32}} were only partially cleared
(32 of 128 bytes for 8 elements);
- non-packed structs were assumed to take 8 bytes per field.
Use LLVM.storage_size with the module's data layout instead.
hipcompile looked for late global hostcalls and externally initialized globals in meta.ir, and read the entry point's name, after the JuliaContext() block had returned, i.e. after the thread-safe context the module belongs to had been disposed of. That relied on the context leaking (GPUCompiler.jl#970): memcheck reports it as a use of the module after its context's disposal, and it would stop working once that leak is fixed. Do this inside the block. The caller owns the module GPUCompiler returns, so also dispose of it there, like GPUCompiler's own callers of compile do; otherwise the module leaks along with the context. The codegen test that calls compile directly does the same.
finish_ir! created a TargetMachine for every kernel it compiled to run the amdgpu-attributor pass with, and never disposed of it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports AMDGPU.jl to LLVM.jl 10.
Main changes:
memcpy!/memset!, the exception CAS) is now an@llvmgeneratedfunction. The unusedget_global_pointerandstring_lengthwere removed instead of ported.Intrinsicdeclarations. The fence overlays usefence!instead of IR strings, and device strings go directly into the global address space.Fixes, each in its own commit:
@ROCStaticLocalArray(Int32, 256)cleared 4096 bytes instead of 1024. The size now comes from the data layout.hipcompileinspected the compiled IR after theJuliaContextthat owns it was gone. It now inspects and disposes of it inside the block.finish_ir!leaked aTargetMachinefor every kernel.The LLVM IR and GCN of 12 kernels covering every ported generator are identical to main's on gfx1030 and gfx90a, apart from the intended fixes. The test suite only ran on a gfx1036 iGPU, which is unsupported and flaky. There, the port does no worse than main (15 failures and 1 error on Julia 1.12, against 187 failures and 111 errors for main). A CI run on a supported GPU is still needed.
Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).
Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.