Skip to content

Implement the atomic functions with UnsafeAtomics - #3350

Open
maleadt wants to merge 5 commits into
mainfrom
tb/unsafeatomics
Open

maleadt wants to merge 5 commits into
mainfrom
tb/unsafeatomics

Conversation

@maleadt

@maleadt maleadt commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Implement the device atomics with UnsafeAtomics.jl, as Metal.jl did in JuliaGPU/Metal.jl#977. UnsafeAtomics emits LLVM atomics with an ordering and a synchronization scope, GPUCompiler (2.12) renames the scope for PTX, and the NVPTX back-end lowers them, expanding what PTX lacks into compare-and-swap loops. The public API doesn't change: the functions keep CUDA C's names (atomic_add! for atomicAdd, threadfence for __threadfence, ...) and their scope argument.

That removes the IR generators, the bitcasts for floating-point compare-and-swap, the inline assembly in grid synchronization, and llvm.nvvm.membar (the fences are LLVM fences now). Block- and system-scope atomic_inc!/atomic_dec! use the native instruction instead of a compare-and-swap loop (from LLVM 16; Julia 1.10 keeps NVVM's intrinsics for device scope).

Technically breaking: the atomic functions are now relaxed, like CUDA C's. They were documented as acquire/release, but the back-ends before LLVM 23 dropped the ordering, so they were effectively relaxed until CUDA.jl 6.4, where they became a full barrier (MEMBAR + CCTL.IVALL in SASS, and no RED for unused results). On an RTX 5080 the difference is small (≤1% on an atomic histogram, ~3% on a scatter kernel with stores in flight), but relaxed is what CUDA C, HIP and Kokkos do. Code that synchronizes through atomics needs a fence or an ordered atomic. The docs now explain that with UnsafeAtomics, which is also the interface for atomic loads and stores, and they recommend Atomix.jl's @atomic over CUDA.@atomic.

Other changes:

  • CUDA.@atomic with an operation that has no native instruction could return without updating the element when another thread stored a NaN with another payload: the loop compared values with isequal. Now it uses the compare-and-swap's success flag (new test).
  • 16-bit compare-and-swap keeps using the native instruction from sm_70, and so does the compare-and-swap loop for 16-bit types (e.g. BFloat16 addition below sm_90). LLVM implements it on the containing 32-bit word, which compute-sanitizer reports as an out-of-bounds read on a 1-element array.
  • The exception output lock now publishes its owner before other threads compare against it, and uses atomic accesses where threads race.

The atomic functions no longer check what the target supports themselves. GPUCompiler now validates atomics in the IR for the PTX target (JuliaGPU/GPUCompiler.jl#996), so atomics used without the atomic functions, e.g. through UnsafeAtomics, Atomix or Enzyme, get the same errors, pointing at the Julia code that performs them. CUDA.jl tells it which platforms lack system-scope atomics: Pascal GPUs under Windows (#3187), and Tegra GPUs before sm_72, which the docstrings mentioned but nothing checked. The messages change, e.g. system-scope atomic operation (requires compute capability 6.0; use device scope if system-wide atomicity is not required).

The ordered atomics in grid synchronization rely on the NVPTX back-end before sm_70 only as far as NVPTX_LLVM_Backend_jll 23.1.2+0 supports (JuliaPackaging/Yggdrasil#15013 fixed the rest in +1), since compat bounds can't require a build.

Before merging: drop the "Dev GPUCompiler." commit and require the GPUCompiler release with the atomics validation in CUDACore's compat ([sources] doesn't apply on Julia 1.10, so that CI job needs the release).

This also unblocks Atomix 1.6 (JuliaConcurrent/Atomix.jl#87), which drops its CUDA extension for CUDACore 6.5.

Tested on an RTX 5080 with Julia 1.10 and 1.12: the device intrinsics, cooperative groups, execution, exceptions, KernelAbstractions, sorting and codegen tests, and the atomics tests under compute-sanitizer. The atomics validation was tested on an RTX 5080 with Julia 1.13: the device, execution, exceptions, KernelAbstractions and codegen tests, except for "device code survives compile-all", which already failed before it (the host cannot compile the inline assembly of the 16-bit compare-and-swap).

UnsafeAtomics emits LLVM atomics with an ordering and a scope, which the NVPTX
back-end lowers, expanding what PTX lacks into compare-and-swap loops. That
replaces the IR generators, the bitcasts for floating-point compare-and-swap,
and the compare-and-swap loop of the generic fallback, which compared values
with isequal and so could give up on an update when it saw another NaN. Fences
become LLVM fences too.

16-bit compare-and-swap keeps using the native instruction from sm_70: LLVM
implements it on the containing 32-bit word, which accesses memory past the
end of an array (as compute-sanitizer reports). Below LLVM 16, which can't
express uinc_wrap and udec_wrap, device-scope atomic_inc! and atomic_dec! keep
using NVVM's intrinsics.
Grid synchronization used inline assembly for its release update and acquire
loads, and volatile loads before sm_70. The exception output lock and status
were accessed with plain loads and stores while other threads updated them
atomically. Use atomics with explicit orderings instead, and publish the owner
of the output lock before other threads compare their index with it.
CUDA C's atomicAdd and friends are relaxed: they're atomic, but don't order
the memory accesses around them. CUDA.jl's were documented as acquire/release,
but back-ends before LLVM 23 dropped the ordering, so in practice they were
relaxed until CUDA.jl 6.4, which made them a full barrier. This is technically
breaking: code that synchronizes through them needs a fence or an ordered
atomic, which the documentation now explains with UnsafeAtomics. It also
recommends Atomix's @atomic over CUDA.@atomic.
GPUCompiler now validates the atomics in the IR for the PTX target, which also covers
atomics used without the atomic functions, e.g. through UnsafeAtomics or Atomix, so the
static assertions in the atomic functions can go. Tell it which platforms lack system-scope
atomics: Pascal GPUs under Windows, and Tegra GPUs before sm_72, which the atomic functions
documented but did not check.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant