Skip to content

Emit atomics as LLVM atomics, lowered by GPUCompiler - #977

Open
maleadt wants to merge 14 commits into
mainfrom
tb/atomics
Open

maleadt wants to merge 14 commits into
mainfrom
tb/atomics

Conversation

@maleadt

@maleadt maleadt commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Depends on JuliaGPU/GPUCompiler.jl#942, which lowers LLVM atomics for Metal (GPUCompiler 2.10), and requires GPUCompiler 2.12 and UnsafeAtomics 0.4 (unreleased):

Metal.jl's atomic functions now emit plain LLVM atomics through UnsafeAtomics, with an ordering and a synchronization scope (device scope for device memory, workgroup scope for threadgroup memory, as in MSL). GPUCompiler then does what an LLVM back-end would: it selects AIR's atomic intrinsics, expands what AIR lacks, and implements orderings with fences before MSL 4.1. That brings:

  • ordered atomics on MSL 3.2 and 4.0, not just 4.1;
  • threadgroup Float32 add/sub on every MSL version;
  • ordered 64-bit min/max.

Other changes:

  • Relaxed loads of device memory keep MSL's semantics: they're only guaranteed to eventually observe stores from the same threadgroup. GPUCompiler implements LLVM's stronger guarantee for device-scope relaxed loads with acquire loads, which is up to 6–7× slower when the loads hit the cache. So Metal.jl emits its relaxed loads at workgroup scope, which compiles exactly as MSL's. To wait for another threadgroup, use an acquire load. The docstring says so, and new tests check that both an acquire load and an UnsafeAtomics device-scope relaxed load see another threadgroup's store.
  • Memory orders have to be compile-time constants, as on main and in MSL (whose headers require it with METAL_CONST_ARG, and whose AIR intrinsics take the order as an immediate). An order that isn't a constant fails to compile.
  • Loads and stores only take the part of an order that applies to them, like Clang does.
  • Memory flags are an optional last argument of the same functions, and are emitted as LLVM atomics too. LLVM's orderings order all memory, so the flags are part of the synchronization scope's name, e.g. device-mem-global for MemoryFlagDevice, which GPUCompiler turns into the flags operand (Rename synchronization scopes for the target GPUCompiler.jl#979). Like the order, the flags have to be compile-time constants. Atomics with flags therefore work from MSL 3.2 instead of 4.1, including Float32 atomics on threadgroup memory. atomic_thread_fence also emits an LLVM fence with the flags in its scope, so acquire and release fences work from MSL 3.2 too (as sequentially-consistent fences before 4.1). Relaxed fences still call air.atomic.fence, as LLVM has no relaxed fences.
  • No direct air.atomic.* calls are left, and GPUCompiler legalizes the ABI of those that other code emits for older targets, so the downgrading code in finish_ir! is gone.
  • The functions on LLVM atomics are generic methods instead of being generated per type and address space, and one compare-exchange loop implements atomic_fetch_op_explicit, with and without flags.
  • Threadgroup arrays and device buffers are padded to whole 4-byte words, since GPUCompiler implements 8- and 16-bit atomics on the containing 32-bit word, and Metal's shader validation drops accesses past a buffer's length.
  • The atomic functions are documented, in a new "Atomics" section of the kernel programming API reference.

New tests cover the LLVM atomics on MSL 3.2, 4.0 and 4.1: contention, partword neighbours, the compare-exchange expansions, acquire/release message passing, waiting on another threadgroup, and the kernel from JuliaGPU/GPUCompiler.jl#934. Checks of the AIR that GPUCompiler selects are left to GPUCompiler's tests; only the intrinsics Metal.jl calls itself are checked here.

Tested on an M1 with GPUCompiler 2.11.1, LLVM.jl 10.0.0 and UnsafeAtomics 0.3.3, before the last two commits:

  • device/intrinsics/atomics, device/intrinsics/synchronization and kernelabstractions, on Julia 1.10 and 1.12: 2439 pass, 8 broken.
  • The full suite on Julia 1.12: 7077 pass, 56 broken. examples/gtk failed because its test worker died, which happens when it runs on its own too.
  • With --validate (Metal API and shader validation): device/intrinsics/atomics, synchronization, array, pool and kernelabstractions, 3087 pass, 8 broken.

The last two commits (method tables, memory flags in scopes) were tested on the M1 on Julia 1.12, with GPUCompiler 2.12.0 and UnsafeAtomics 0.4: device/intrinsics/atomics (229 pass) and device/intrinsics/synchronization (19 pass).

Comment thread src/device/intrinsics/atomics.jl Outdated
@codecov

codecov Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.95918% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 87.28%. Comparing base (c51cdc2) to head (33ee27e).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/device/intrinsics/atomics.jl 96.66% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #977      +/-   ##
==========================================
- Coverage   87.29%   87.28%   -0.02%     
==========================================
  Files          92       92              
  Lines        6816     6770      -46     
==========================================
- Hits         5950     5909      -41     
+ Misses        866      861       -5     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Metal Benchmarks

Details
Benchmark suite Current: 33ee27e Previous: c51cdc2 Ratio
array/accumulate/Float32/1d 391583 ns 393209 ns 1.00
array/accumulate/Float32/dims=1 366458 ns 370125 ns 0.99
array/accumulate/Float32/dims=1L 8888000 ns 8891250 ns 1.00
array/accumulate/Float32/dims=2 430000 ns 437125 ns 0.98
array/accumulate/Float32/dims=2L 2577250 ns 2721125 ns 0.95
array/accumulate/Int64/1d 839958 ns 848000 ns 0.99
array/accumulate/Int64/dims=1 913375 ns 907708 ns 1.01
array/accumulate/Int64/dims=1L 9547791 ns 9559042 ns 1.00
array/accumulate/Int64/dims=2 1215666 ns 1205958 ns 1.01
array/accumulate/Int64/dims=2L 6566125 ns 6545792 ns 1.00
array/broadcast 239583 ns 240208 ns 1.00
array/construct 2250 ns 2375 ns 0.95
array/permutedims/2d 457250 ns 458750 ns 1.00
array/permutedims/3d 1031625 ns 1033750 ns 1.00
array/permutedims/4d 1115166 ns 1130792 ns 0.99
array/private/copy 226000 ns 224541 ns 1.01
array/private/copyto!/cpu_to_gpu 212708 ns 214917 ns 0.99
array/private/copyto!/gpu_to_cpu 215292 ns 216250 ns 1.00
array/private/copyto!/gpu_to_gpu 221625 ns 220167 ns 1.01
array/private/iteration/findall/bool 1055542 ns 1057500 ns 1.00
array/private/iteration/findall/int 1216709 ns 1212709 ns 1.00
array/private/iteration/findfirst/bool 971125 ns 972583 ns 1.00
array/private/iteration/findfirst/int 1024541 ns 1020583 ns 1.00
array/private/iteration/findmin/1d 1086750 ns 1088625 ns 1.00
array/private/iteration/findmin/2d 1019875 ns 1021375 ns 1.00
array/private/iteration/logical 1546375 ns 1547084 ns 1.00
array/private/iteration/scalar 1400625 ns 1403375 ns 1.00
array/random/rand/Float32 419917 ns 417583 ns 1.01
array/random/rand/Int64 498500 ns 500375 ns 1.00
array/random/rand!/Float32 404750 ns 403625 ns 1.00
array/random/rand!/Int64 430292 ns 426250 ns 1.01
array/random/randn/Float32 390833 ns 390584 ns 1.00
array/random/randn!/Float32 368583 ns 375375 ns 0.98
array/reductions/mapreduce/Float32/1d 272458 ns 274667 ns 0.99
array/reductions/mapreduce/Float32/dims=1 350667 ns 354750 ns 0.99
array/reductions/mapreduce/Float32/dims=1L 608417 ns 609625 ns 1.00
array/reductions/mapreduce/Float32/dims=2 356375 ns 354292 ns 1.01
array/reductions/mapreduce/Float32/dims=2L 1241417 ns 1240917 ns 1.00
array/reductions/mapreduce/Int64/1d 435833 ns 455583 ns 0.96
array/reductions/mapreduce/Int64/dims=1 632958 ns 635500 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 1011959 ns 1015750 ns 1.00
array/reductions/mapreduce/Int64/dims=2 789792 ns 787291 ns 1.00
array/reductions/mapreduce/Int64/dims=2L 2196750 ns 2191209 ns 1.00
array/reductions/reduce/Float32/1d 271792 ns 274084 ns 0.99
array/reductions/reduce/Float32/dims=1 358084 ns 355625 ns 1.01
array/reductions/reduce/Float32/dims=1L 617000 ns 620375 ns 0.99
array/reductions/reduce/Float32/dims=2 250792 ns 251208 ns 1.00
array/reductions/reduce/Float32/dims=2L 477959 ns 471792 ns 1.01
array/reductions/reduce/Int64/1d 457125 ns 458333 ns 1.00
array/reductions/reduce/Int64/dims=1 637584 ns 641333 ns 0.99
array/reductions/reduce/Int64/dims=1L 1000042 ns 1010167 ns 0.99
array/reductions/reduce/Int64/dims=2 260084 ns 257625 ns 1.01
array/reductions/reduce/Int64/dims=2L 664709 ns 670834 ns 0.99
array/shared/copy 128416 ns 130250 ns 0.99
array/shared/copyto!/cpu_to_gpu 37791 ns 37291 ns 1.01
array/shared/copyto!/gpu_to_cpu 36750 ns 37875 ns 0.97
array/shared/copyto!/gpu_to_gpu 37083 ns 38209 ns 0.97
array/shared/iteration/findall/bool 1059042 ns 1048458 ns 1.01
array/shared/iteration/findall/int 1218917 ns 1223209 ns 1.00
array/shared/iteration/findfirst/bool 980667 ns 976875 ns 1.00
array/shared/iteration/findfirst/int 991125 ns 981250 ns 1.01
array/shared/iteration/findmin/1d 1049625 ns 1050875 ns 1.00
array/shared/iteration/findmin/2d 1019834 ns 1027416 ns 0.99
array/shared/iteration/logical 1522500 ns 1525334 ns 1.00
array/shared/iteration/scalar 76.375 ns 75.89014373716633 ns 1.01
array/sorting/1d 1913500 ns 1921125 ns 1.00
array/sorting/2d 8265125 ns 8319917 ns 0.99
integration/byval/reference 1116125 ns 1120625 ns 1.00
integration/byval/slices=1 1124542 ns 1122291 ns 1.00
integration/byval/slices=2 2031375 ns 2029542 ns 1.00
integration/byval/slices=3 6856667 ns 6806458 ns 1.01
integration/metaldevrt 395125 ns 396167 ns 1.00
kernel/indexing 204708 ns 207625 ns 0.99
kernel/indexing_checked 389250 ns 395208 ns 0.98
kernel/launch 2051 ns 2101.8888888888887 ns 0.98
kernel/rand 406708 ns 408958 ns 0.99
latency/import 1863992125 ns 1862270958 ns 1.00
latency/precompile 33417766792 ns 31631620958 ns 1.06
latency/ttfp 2293500584 ns 2295776542 ns 1.00
metal/synchronization/context 529.3842105263158 ns 525.958115183246 ns 1.01
metal/synchronization/stream 629.3604651162791 ns 657.375796178344 ns 0.96

This comment was automatically generated by workflow using github-action-benchmark.

maleadt added 13 commits October 5, 2026 21:05
GPUCompiler implements 8- and 16-bit atomics as operations on the
containing 32-bit word, like LLVM's AtomicExpand, which assumes that the
whole word can be accessed. That holds for device buffers, which Metal
allocates in pages, so make it hold for threadgroup arrays too.
GPUCompiler implements 8- and 16-bit atomics on the containing 32-bit
word. That word always exists in memory, as Metal allocates buffers in
pages, but Metal's shader validation only allows accesses within the
buffer's length, and drops the others: a UInt16 atomic on a 2-byte
array had no effect. Pad the buffers that aren't backed by host memory
to whole words, as threadgroup arrays are.
GPUCompiler now lowers LLVM atomics for Metal (JuliaGPU/GPUCompiler.jl#942),
like an LLVM back-end would: selecting AIR's atomic intrinsics, expanding
what AIR lacks (8- and 16-bit operations, compare-exchange loops for
operations like nand or floating-point min/max), and implementing
orderings with fences before MSL 4.1. Emit the atomic functions as LLVM
atomics through UnsafeAtomics, with the device or workgroup scope for
device or threadgroup memory, as MSL does.

This supports ordered atomics on MSL 3.2 and 4.0, threadgroup Float32
add/sub on every version, and ordered 64-bit min/max. Orders passed as
values are dispatched on, so that inference stays precise. Loads and
stores only take the part of an order they can have, like Clang. Calls
with explicit memory flags still use the MSL 4.1 intrinsics, whose ABI
GPUCompiler now legalizes for older targets, so Metal.jl no longer does.

This needs GPUCompiler 2.10 for the lowering; 2.11 is the first release
that supports LLVM.jl 10.
GPUCompiler now sets it on every atomic load and read-modify-write it
selects (without it, Apple's compiler turns read-modify-writes that don't
change memory into loads that can be hoisted out of loops). Intrinsics
called with explicit memory flags keep the volatile bit they are given.
MSL only makes device memory coherent within a threadgroup by default,
so its relaxed loads are only guaranteed to observe stores from the same
threadgroup; on an M1, a loop relaxed-loading a flag that another
threadgroup sets never observes the store. LLVM's device-scope relaxed
loads have to observe every thread's stores, which GPUCompiler now
implements by making the ones that may be repeated acquire loads. Keep
`atomic_load_explicit` with the relaxed order at MSL's semantics (and
cost) by emitting it with the workgroup scope, and document how to wait
for another threadgroup: with an acquire load, or a device-scope relaxed
load through UnsafeAtomics. Test both, with bounded waits.
- The functions on LLVM atomics are generic over the element type and
  address space, instead of generated per combination.
- `with_llvm_order` branches on the order once, for values and `Val`s.
- The functions with memory flags come from one table, and call the
  Float32 intrinsics MSL uses instead of reinterpreting through UInt32.
  Like MSL, they then require Metal 4.1 for Float32 atomics on
  threadgroup memory (the functions without flags support those on
  every version).
- One compare-exchange loop implements `atomic_fetch_op_explicit`, with
  and without memory flags.
GPUCompiler tests the AIR it selects for LLVM atomics, and how it
legalizes the MSL 4.1 intrinsics, for every target. Only check the
intrinsics Metal.jl calls itself, and share the list of targets.
MSL's atomic functions require the memory order to be a compile-time
constant: its headers check this with METAL_CONST_ARG, and the AIR
intrinsics take the order as an immediate. So did Metal.jl before this
PR, by converting the order to a Val. Passing an order as a run-time
value only worked because `with_llvm_order` branched on it.

Require constants again. `llvm_order` maps a constant `memory_order`
to an UnsafeAtomics ordering, and an order that isn't a constant is a
dynamic call, i.e. an InvalidIRError. The default orders are Vals, so
they don't depend on the functions being inlined, which GPUCompiler's
interpreter doesn't always do.
On Julia 1.11, calling `fill` while generating the atomic functions
resolves Metal's `fill` binding to Base's, so that defining it later in
array.jl fails to precompile.
On Julia 1.11, GPUCompiler rejects the boxed constant that a run-time
order needs with an ErrorException, before validating the IR.
GPUCompiler then stacks its shared method table below Metal's tables,
whose overlays apply to code for every back-end.
GPUCompiler reads MSL's memory flags from the synchronization scope of
an LLVM atomic or fence (e.g., device-mem-global), as LLVM's orderings
order all memory. Emit the atomic functions and fences that take flags
as LLVM atomics and fences too, so that they are lowered like the
others, instead of calling the MSL 4.1 intrinsics. They then work on
every target with ordered atomics (MSL 3.2), and threadgroup Float32
atomics take flags on every version. Relaxed fences, which LLVM cannot
express, still call the intrinsic.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants