Skip to content

Generate inline asm and IR helpers with LLVM.jl 10's @llvmgenerated - #175

Draft
AntonOresten wants to merge 1 commit into
mainfrom
agent/llvmgenerated
Draft

AntonOresten wants to merge 1 commit into
mainfrom
agent/llvmgenerated

Conversation

@AntonOresten

Copy link
Copy Markdown
Member

Switches PTX's hand-written LLVM IR strings to LLVM.jl 10's generate_llvmcall / @llvmgenerated (JuliaLLVM/LLVM.jl#596; cf. JuliaGPU/CUDA.jl#3305).

  • Inline asm: convergent_asmcall / plain_asmcall build the asm call with an IRBuilder and attach the call-site attributes @asmcall can't (convergent nomerge nounwind). LLVM.jl derives the signature from the Julia types, handles typed (Julia 1.10/1.11) and opaque pointers, and verifies the IR at generation time. Wrappers get back the call expression and splice it into @eval, replacing the *_asm_ir string builders.
  • @llvmgenerated helpers: reinterpret_addrspace, vector global ld/st (_vec_load/_vec_store), and the mma half-vector bitcasts.
  • Compat: LLVM = "10". That pulls in CUDA.jl ≥ 6.4.2 / GPUCompiler ≥ 2.11.
  • Tests: the tcgen05 ld/st and mbarrier tests check the IR in the wrappers' typed code instead of calling internal string builders (llvmcall_ir in test/setup.jl).

Net −115 lines.

Precompile time

This is the open question. Base.compilecache(PTX), 3 runs each on LLVM.jl 10.0.0 / Julia 1.13:

main 19.3–20.3 s
this PR 21.8–23.1 s

Most of the gap comes from generate_llvmcall assigning every argument expression to a temporary before the llvmcall. PTX expands thousands of these, many with 100+ arguments (wgmma/mma), and lowering all those assignments is the cost. When I stripped the temporaries as an experiment, the module body matched main. The temporaries are only needed when an argument is dropped (static/singleton): when every argument is passed through, llvmcall already evaluates them once, in order. I plan to propose that change upstream. Specializations of the asm helpers on hundreds of distinct tuple types were a ~3 s regression of their own; that's already fixed here.

Testing

  • Julia 1.13 full suite on GB10: passes except host/conformance. That failure is unrelated: a fresh resolve picks NVPTX_LLVM_Backend_jll 23.1.2, but the registry was generated against 23.1.1, so main hits it too.
  • Julia 1.10, typed-pointer path: host, ptxas and GPU subset (wrappers, address_roles, vector_results, tcgen05_ldst, mbarrier_forms, inst, tma, aqua, wait_registers, spcompress, ptxas baseline/golden/tma, rms_norm, hopper gemm_warpgroup, ampere and blackwell flash_attention): all pass.

🤖 Generated with Claude Code

Build the convergent and plain inline-asm calls with `generate_llvmcall`
instead of hand-written IR strings: the builder attaches the call-site
attributes `@asmcall` cannot (`convergent nomerge nounwind`), LLVM.jl
derives the signature, handles typed and opaque pointers, and verifies the
IR. Wrappers now call `convergent_asmcall`/`plain_asmcall` directly and get
back the call expression.

`reinterpret_addrspace`, the vector global loads/stores and the mma
half-vector bitcasts become `@llvmgenerated` functions.

Requires LLVM.jl 10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant