Repository navigation
Conversation
Import the IR and Build vocabularies next to LLVM.Interop, since `using LLVM` no longer exports the API, and use properties instead of the removed accessor functions (metadata, function attributes, module functions, linkage, initializer, alignment, the entry block and the subprogram of a function, ...). create_function and call_function were removed: the additional kernel arguments, the constant tables for randn and the invariant loads now generate their IR with generate_llvmcall. Insertion points are explicit now: the block that initializes the RNG state is created LLVM.before the entry block, and the builder positioned LLVM.at_end of it. The builder gets the line-0 location of the kernel's subprogram for the instructions it inserts; the old `debuglocation!(builder, first(instructions(top_bb)))` copied that location onto the first instruction of the original entry block (or cleared its location without a subprogram), which wasn't the intent.
Look up or declare functions with get! on the module's function view, check for the RNG attribute by its kind instead of collecting and comparing attributes, use the scoped linkage enumeration instead of LLVM.API, and the public metadata kinds. The intrinsics for the additional kernel arguments now only get their readnone attribute when they're declared, rather than every time they are looked up.
The accessor for the additional kernel arguments, the invariant load and the constant tables for randn were @generated functions around generate_llvmcall: define them with @llvmgenerated instead, which derives the LLVM signature from the Julia one. The invariant load's checks move to a wrapper, and its generator uses an unconditional typed bitcast (which folds away with opaque pointers) and load!'s align keyword. The table accessors become plain functions that pass the table's name and element type to one generator, which reads the (constant) table from Random. Use memory effects instead of the readnone attribute, which is invalid on functions as of LLVM 16, derive the return type of the deferred_codegen declaration from Julia's lowering of Ptr{Cvoid} instead of a version decision tree, and implement NDIteration.assume with LLVM.Interop.assume instead of a copy of its IR.
GPUCompiler now hands the module that Julia's code generator produced to the caller of compile: on Julia 1.11 and later it moves the module out of the native code, and on 1.10 (LLVM 15, which can't do that) it returns the module for the caller to consume or dispose of. compile_to_obj only read the entry point's name and attributes, and never disposed of the module. That leaked it: up to Julia 1.13, the native-code descriptor is never freed and keeps its thread-safe module, and with it the context, alive, so the module isn't freed together with the context either. memcheck also reported every compiled module. Dispose of the IR with @dispose inside the JuliaContext block, once the entry point has been inspected, like GPUCompiler's own callers do, so that it is also disposed of when the inspection throws.
The index was decremented in its own type and passed to getelementptr as is, which sign-extends narrower indices, so e.g. a UInt32 index of 0x80000001 addressed a negative offset. Convert it to Int in Julia first, like LLVM.jl's unsafe_load does.
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #823 +/- ##
==========================================
- Coverage 70.77% 67.89% -2.89%
==========================================
Files 27 27
Lines 2221 2327 +106
==========================================
+ Hits 1572 1580 +8
- Misses 649 747 +98 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports KernelAbstractions to LLVM.jl 10. LLVM.jl is only used by the POCL back-end, a copy of OpenCL.jl's compiler and device code, and by one
llvmcallinNDIteration. The port follows OpenCL.jl's (JuliaGPU/OpenCL.jl#525).Main changes:
generate_llvmcallfor the IR generators.@llvmgeneratedfor the kernel-argument, invariant-load andrandntable generators. Memory effects replacereadnone, which is invalid on LLVM 16+, andassumeusesLLVM.Interop.assume.Fixes, each in its own commit:
unsafe_invariant_loadsign-extended narrow indices, so a largeUInt32index addressed memory before the array.Tested on Julia 1.10, 1.12 and 1.13, with the same results as main. LLVM.jl's memcheck mode on 1.12 reports nothing.
Requires GPUCompiler 2.11 (JuliaGPU/GPUCompiler.jl#974), GPUToolbox 3.3.3 (JuliaGPU/GPUToolbox.jl#28), SPIRVIntrinsics 1.2 (JuliaGPU/OpenCL.jl#525), and UnsafeAtomics 0.3.3 (JuliaConcurrent/UnsafeAtomics.jl#31) or 0.4.
Disclaimer: this PR is AI-assisted and has not been reviewed in detail. Tests pass locally, so it should be a good starting point for a maintainer to complete the upgrade.