Repository navigation
Conversation
UnsafeAtomics emits LLVM atomics with an ordering and a scope, which the NVPTX back-end lowers, expanding what PTX lacks into compare-and-swap loops. That replaces the IR generators, the bitcasts for floating-point compare-and-swap, and the compare-and-swap loop of the generic fallback, which compared values with isequal and so could give up on an update when it saw another NaN. Fences become LLVM fences too. 16-bit compare-and-swap keeps using the native instruction from sm_70: LLVM implements it on the containing 32-bit word, which accesses memory past the end of an array (as compute-sanitizer reports). Below LLVM 16, which can't express uinc_wrap and udec_wrap, device-scope atomic_inc! and atomic_dec! keep using NVVM's intrinsics.
Grid synchronization used inline assembly for its release update and acquire loads, and volatile loads before sm_70. The exception output lock and status were accessed with plain loads and stores while other threads updated them atomically. Use atomics with explicit orderings instead, and publish the owner of the output lock before other threads compare their index with it.
CUDA C's atomicAdd and friends are relaxed: they're atomic, but don't order the memory accesses around them. CUDA.jl's were documented as acquire/release, but back-ends before LLVM 23 dropped the ordering, so in practice they were relaxed until CUDA.jl 6.4, which made them a full barrier. This is technically breaking: code that synchronizes through them needs a fence or an ordered atomic, which the documentation now explains with UnsafeAtomics. It also recommends Atomix's @atomic over CUDA.@atomic.
GPUCompiler now validates the atomics in the IR for the PTX target, which also covers atomics used without the atomic functions, e.g. through UnsafeAtomics or Atomix, so the static assertions in the atomic functions can go. Tell it which platforms lack system-scope atomics: Pascal GPUs under Windows, and Tegra GPUs before sm_72, which the atomic functions documented but did not check.
This was referenced Oct 6, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implement the device atomics with UnsafeAtomics.jl, as Metal.jl did in JuliaGPU/Metal.jl#977. UnsafeAtomics emits LLVM atomics with an ordering and a synchronization scope, GPUCompiler (2.12) renames the scope for PTX, and the NVPTX back-end lowers them, expanding what PTX lacks into compare-and-swap loops. The public API doesn't change: the functions keep CUDA C's names (
atomic_add!foratomicAdd,threadfencefor__threadfence, ...) and their scope argument.That removes the IR generators, the bitcasts for floating-point compare-and-swap, the inline assembly in grid synchronization, and
llvm.nvvm.membar(the fences are LLVM fences now). Block- and system-scopeatomic_inc!/atomic_dec!use the native instruction instead of a compare-and-swap loop (from LLVM 16; Julia 1.10 keeps NVVM's intrinsics for device scope).Technically breaking: the atomic functions are now relaxed, like CUDA C's. They were documented as acquire/release, but the back-ends before LLVM 23 dropped the ordering, so they were effectively relaxed until CUDA.jl 6.4, where they became a full barrier (
MEMBAR+CCTL.IVALLin SASS, and noREDfor unused results). On an RTX 5080 the difference is small (≤1% on an atomic histogram, ~3% on a scatter kernel with stores in flight), but relaxed is what CUDA C, HIP and Kokkos do. Code that synchronizes through atomics needs a fence or an ordered atomic. The docs now explain that with UnsafeAtomics, which is also the interface for atomic loads and stores, and they recommend Atomix.jl's@atomicoverCUDA.@atomic.Other changes:
CUDA.@atomicwith an operation that has no native instruction could return without updating the element when another thread stored a NaN with another payload: the loop compared values withisequal. Now it uses the compare-and-swap's success flag (new test).The atomic functions no longer check what the target supports themselves. GPUCompiler now validates atomics in the IR for the PTX target (JuliaGPU/GPUCompiler.jl#996), so atomics used without the atomic functions, e.g. through UnsafeAtomics, Atomix or Enzyme, get the same errors, pointing at the Julia code that performs them. CUDA.jl tells it which platforms lack system-scope atomics: Pascal GPUs under Windows (#3187), and Tegra GPUs before sm_72, which the docstrings mentioned but nothing checked. The messages change, e.g.
system-scope atomic operation (requires compute capability 6.0; use device scope if system-wide atomicity is not required).The ordered atomics in grid synchronization rely on the NVPTX back-end before sm_70 only as far as
NVPTX_LLVM_Backend_jll23.1.2+0 supports (JuliaPackaging/Yggdrasil#15013 fixed the rest in +1), since compat bounds can't require a build.Before merging: drop the "Dev GPUCompiler." commit and require the GPUCompiler release with the atomics validation in CUDACore's compat (
[sources]doesn't apply on Julia 1.10, so that CI job needs the release).This also unblocks Atomix 1.6 (JuliaConcurrent/Atomix.jl#87), which drops its CUDA extension for CUDACore 6.5.
Tested on an RTX 5080 with Julia 1.10 and 1.12: the device intrinsics, cooperative groups, execution, exceptions, KernelAbstractions, sorting and codegen tests, and the atomics tests under compute-sanitizer. The atomics validation was tested on an RTX 5080 with Julia 1.13: the device, execution, exceptions, KernelAbstractions and codegen tests, except for "device code survives compile-all", which already failed before it (the host cannot compile the inline assembly of the 16-bit compare-and-swap).