Conversation
The indexing, shared memory, malloc, atomics, assertion, printf, ldg and string pointer intrinsics, and the constant tables for randn, all built their IR with create_function/call_function around a lot of context and builder boilerplate. LLVM.jl 9.14's @llvmgenerated and generate_llvmcall derive the LLVM signature from the Julia one and verify the generated IR, so use those instead. Julia code that ran before the IR, like checking the atomic scope or adjusting the index for ldg, moves to a wrapper that calls the @llvmgenerated function. The @llvmgenerated helper for ldg is defined on every LLVM version, not only on LLVM 20 and later where it is used: the device function test in core/device/method_table collects every @device_function definition regardless of @static conditions, and failed on older LLVM because the helper didn't exist there.
Bump the LLVM.jl compat of CUDACore, CUDATools and cuBLAS to 10, and import the IR and Build vocabularies in CUDACore now that `using LLVM` only exports @dispose. Accessors and setters become properties, the builder and new block are positioned with insertion points, named metadata is created explicitly with get!, and the empty module pass manager that was only run to attach that metadata is removed.
Refer to linkages, atomic orderings, atomicrmw operations and the code generation file type through LLVM.jl's scoped enumerations instead of LLVM.API, and look up atomicrmw operations by their IR name rather than through a hand-written table. Look up the target globals with get, and get or declare deferred_codegen with get!.
The cluster barriers were IR strings that declared the intrinsic with the attributes older LLVM can't infer, and the volatile load in the grid barrier was an IR string with a branch for typed pointers. Both are built with @llvmgenerated, declaring the intrinsic and its attributes with LLVM.jl, which handles the pointer representation on every LLVM version. The same goes for the weak shared-memory globals that hold the RNG state, which keep their names and linkage, and for the arithmetic with contract and nsz flags in the volumerhs benchmark, which interpolated the operation and type into IR. The constant tables of randn were generated by a closure that captured the table and its element type, which generate_llvmcall compiles again for every table. Use a named generator that takes the table's name and element type as statically-known arguments instead, so that it is compiled once. The assertion generator named its strings with a global counter, which made the generated IR depend on session state. The strings are private globals, whose names LLVM keeps unique, so give them fixed names.
Declare the indexing intrinsics with Intrinsic, which gives them LLVM's
declaration and attributes, looking them up with tryparse(Intrinsic,
name) and falling back to a declaration by name for the cluster
registers that LLVM only knows since version 17. Convert between
pointers and Julia's representation of them with pointercast! rather
than ptrtoint!, and derive the type deferred_codegen returns from
Ptr{Cvoid} instead of spelling out its lowering per version.
Poll the grid barrier with LLVM.Interop's volatile_load instead of a
generator of our own, pass the atomics' synchronization scope by name so
that it resolves in the builder's context, pass the alignment of the
ldg load to load! and refer to the metadata kinds that are part of
LLVM.IR without qualification.
The initialization code that finish_module! prepends to kernels using the RNG called `debuglocation!(builder, inst)`, meaning to take the location of the kernel's first instruction, but that copied the builder's location to the instruction instead: with debug info the first instruction got a line-zero location, and without it lost its location. Give the generated instructions the line-zero location, and leave the kernel's own instructions alone.
compile() returned the IR from the JuliaContext block and then looked at its functions and at the entry point's name after the block had returned. That relied on the context leaking (GPUCompiler.jl#970): the IR is only meant to be used while its context is, and memcheck reports using it afterwards. Inspect the IR inside the block, and dispose of it there now that nothing else needs it, so that it doesn't leak the module and keeps working once that leak is fixed.
This was referenced Oct 3, 2026
Contributor
CUDA.jl BenchmarksDetails
This comment was automatically generated by workflow using github-action-benchmark. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports CUDA.jl to LLVM.jl 10. Supersedes #3305, whose commit comes first here.
Main changes:
link_libraries!unnecessary.@llvmgeneratedfor the remaining IR string templates (cluster barriers, grid barrier, RNG state, volumerhs benchmark). The indexing intrinsics are declared withIntrinsic, pointers are converted withpointercast!, the grid barrier usesvolatile_load, and atomics pass their syncscope by name. The generated IR is unchanged.Fixes and behavior changes:
compile()inspected the compiled IR after theJuliaContextthat owns it was gone. It now inspects and disposes of the IR inside the block._pointerref_ldgfrom Generate IR with LLVM.jl's @llvmgenerated and generate_llvmcall #3305 is now defined on every LLVM version, which fixes the device-function test on older LLVM.mainbranch.Tested on an RTX 5080: the full suite on Julia 1.12 and the core tests on 1.10, 1.11 and 1.13 match main. There is one timing-dependent
unsafe_wrapfailure incore/array.jl, which main also has. LLVM.jl's memcheck mode on 1.12 reports nothing.Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).
Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.