Repository navigation
Conversation
Local memory, invariant loads, the constant tables for randn, printf and the additional kernel arguments all built their IR with create_function/call_function around a lot of context and builder boilerplate. LLVM.jl 9.14's @llvmgenerated and generate_llvmcall derive the LLVM signature from the Julia one and verify the generated IR, so use those instead. printf needs the Julia types of its varargs, so it stays a @generated function that uses generate_llvmcall.
Import the IR and Build vocabularies next to LLVM.Interop, since `using LLVM` no longer exports the API, and use properties instead of the removed accessor functions (linkage, initializer, alignment, metadata, function attributes, module functions and globals, ...). Insertion points are explicit now: the block that initializes the RNG state is created with `LLVM.before` the entry block, and the builder positioned `LLVM.at_end` of it. That prologue gets its debug location from the builder only. The old `debuglocation!(builder, first(instructions(top_bb)))` copied the builder's location to the first instruction of the original entry block (despite what its docstring said), overwriting that instruction's location with the synthetic line 0, or clearing it when the kernel had no debug info, so it is dropped rather than ported.
Look up or declare functions with `get!` on the module's function view, look up the feature bitset global with `get`, and check for the RNG attribute by its kind instead of collecting and comparing attributes.
The work-item built-ins, the barriers and the feature bitset were fixed IR strings that declare a global or function with attributes and load from or call it, and the accessor for the additional kernel arguments was a hand-written @generated function. Build them with @llvmgenerated instead, and declare the built-ins and the additional argument intrinsics with memory effects instead of `readnone`. Also make the invariant-load generator use an unconditional typed bitcast (which folds away with opaque pointers), `load!`'s `align` keyword and the public metadata kinds, derive the return type of the `deferred_codegen` declaration from Julia's lowering of `Ptr{Cvoid}` instead of a version decision tree, and drop printf's unused list of argument types. The IR generators of printf and of the RNG's constant tables were closures capturing the format string and the element type, which generate_llvmcall compiles for every one of them. Make them a named @nospecialize function that takes the format string and argument types as statically-known arguments, and a callable object with abstractly typed fields, so that they are compiled once. The generated IR is unchanged.
The index was decremented in its own type and passed to getelementptr as is, which sign-extends narrower indices, so e.g. a UInt32 index of 0x80000001 addressed a negative offset. Convert it to Int in Julia first, like LLVM.jl's unsafe_load does.
GPUCompiler now hands the module that Julia's code generator produced to the caller of compile: on Julia 1.11 and later it moves the module out of the native code, and on 1.10 (LLVM 15, which can't do that) it returns the module for the caller to consume or dispose of. compile_to_obj only read the entry point's name and attributes, and never disposed of the module. That leaked it: up to Julia 1.13, the native-code descriptor is never freed and keeps its thread-safe module, and with it the context, alive, so the module isn't freed together with the context either. memcheck reported every compiled module as leaked. Dispose of the IR with @dispose inside the JuliaContext block, once the entry point has been inspected, so that it is also disposed of when the inspection throws. The test that compiles kernels for a device without Float64 discarded the IR returned by compile(:llvm) as well; dispose of it there too.
This was referenced Oct 3, 2026
SPIRVIntrinsics 1.2.0 requires LLVM.jl 10, so OpenCL.jl, which uses it through lib/intrinsics, requires this version too. SPIRVIntrinsics is released from this repository, and oneAPI and KernelAbstractions need the new version as well.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #525 +/- ##
==========================================
+ Coverage 85.56% 86.29% +0.73%
==========================================
Files 19 19
Lines 1683 1642 -41
==========================================
- Hits 1440 1417 -23
+ Misses 243 225 -18 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports OpenCL.jl and SPIRVIntrinsics to LLVM.jl 10, and releases SPIRVIntrinsics 1.2.0, which oneAPI.jl and KernelAbstractions.jl need. Supersedes #512, whose commit comes first here.
Main changes:
LLVM.IR/LLVM.Buildimports, properties instead of accessors, and explicit insertion points.get/get!on the module's function and global views instead of hand-written get-or-create code.@llvmgeneratedfor the work-item built-ins, barriers, feature bitset and kernel-argument accessor. Memory effects replacereadnone, which is invalid on LLVM 16+. The printf and Ziggurat generators are compiled once instead of per format string or table. The generated IR is unchanged.Fixes, each in its own commit:
UInt32index above 2^31 addressed memory before the array.compile_to_objnever disposed of the IR that GPUCompiler returns, which leaked every compiled module.Tested with PoCL and the Intel CPU runtime on Julia 1.10, 1.12 and 1.13. Results match main, apart from a flaky crash inside
libintelocl.sothat main also has. LLVM.jl's memcheck mode on 1.12 reports nothing.Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) and GPUArrays 11.5.16 (JuliaGPU/GPUArrays.jl#801).
Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.