Conversation
This was referenced Sep 24, 2026
maleadt
marked this pull request as ready for review
October 5, 2026 09:53
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #977 +/- ##
==========================================
- Coverage 87.29% 87.28% -0.02%
==========================================
Files 92 92
Lines 6816 6770 -46
==========================================
- Hits 5950 5909 -41
+ Misses 866 861 -5 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Contributor
There was a problem hiding this comment.
Metal Benchmarks
Details
| Benchmark suite | Current: 33ee27e | Previous: c51cdc2 | Ratio |
|---|---|---|---|
array/accumulate/Float32/1d |
391583 ns |
393209 ns |
1.00 |
array/accumulate/Float32/dims=1 |
366458 ns |
370125 ns |
0.99 |
array/accumulate/Float32/dims=1L |
8888000 ns |
8891250 ns |
1.00 |
array/accumulate/Float32/dims=2 |
430000 ns |
437125 ns |
0.98 |
array/accumulate/Float32/dims=2L |
2577250 ns |
2721125 ns |
0.95 |
array/accumulate/Int64/1d |
839958 ns |
848000 ns |
0.99 |
array/accumulate/Int64/dims=1 |
913375 ns |
907708 ns |
1.01 |
array/accumulate/Int64/dims=1L |
9547791 ns |
9559042 ns |
1.00 |
array/accumulate/Int64/dims=2 |
1215666 ns |
1205958 ns |
1.01 |
array/accumulate/Int64/dims=2L |
6566125 ns |
6545792 ns |
1.00 |
array/broadcast |
239583 ns |
240208 ns |
1.00 |
array/construct |
2250 ns |
2375 ns |
0.95 |
array/permutedims/2d |
457250 ns |
458750 ns |
1.00 |
array/permutedims/3d |
1031625 ns |
1033750 ns |
1.00 |
array/permutedims/4d |
1115166 ns |
1130792 ns |
0.99 |
array/private/copy |
226000 ns |
224541 ns |
1.01 |
array/private/copyto!/cpu_to_gpu |
212708 ns |
214917 ns |
0.99 |
array/private/copyto!/gpu_to_cpu |
215292 ns |
216250 ns |
1.00 |
array/private/copyto!/gpu_to_gpu |
221625 ns |
220167 ns |
1.01 |
array/private/iteration/findall/bool |
1055542 ns |
1057500 ns |
1.00 |
array/private/iteration/findall/int |
1216709 ns |
1212709 ns |
1.00 |
array/private/iteration/findfirst/bool |
971125 ns |
972583 ns |
1.00 |
array/private/iteration/findfirst/int |
1024541 ns |
1020583 ns |
1.00 |
array/private/iteration/findmin/1d |
1086750 ns |
1088625 ns |
1.00 |
array/private/iteration/findmin/2d |
1019875 ns |
1021375 ns |
1.00 |
array/private/iteration/logical |
1546375 ns |
1547084 ns |
1.00 |
array/private/iteration/scalar |
1400625 ns |
1403375 ns |
1.00 |
array/random/rand/Float32 |
419917 ns |
417583 ns |
1.01 |
array/random/rand/Int64 |
498500 ns |
500375 ns |
1.00 |
array/random/rand!/Float32 |
404750 ns |
403625 ns |
1.00 |
array/random/rand!/Int64 |
430292 ns |
426250 ns |
1.01 |
array/random/randn/Float32 |
390833 ns |
390584 ns |
1.00 |
array/random/randn!/Float32 |
368583 ns |
375375 ns |
0.98 |
array/reductions/mapreduce/Float32/1d |
272458 ns |
274667 ns |
0.99 |
array/reductions/mapreduce/Float32/dims=1 |
350667 ns |
354750 ns |
0.99 |
array/reductions/mapreduce/Float32/dims=1L |
608417 ns |
609625 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2 |
356375 ns |
354292 ns |
1.01 |
array/reductions/mapreduce/Float32/dims=2L |
1241417 ns |
1240917 ns |
1.00 |
array/reductions/mapreduce/Int64/1d |
435833 ns |
455583 ns |
0.96 |
array/reductions/mapreduce/Int64/dims=1 |
632958 ns |
635500 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=1L |
1011959 ns |
1015750 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2 |
789792 ns |
787291 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2L |
2196750 ns |
2191209 ns |
1.00 |
array/reductions/reduce/Float32/1d |
271792 ns |
274084 ns |
0.99 |
array/reductions/reduce/Float32/dims=1 |
358084 ns |
355625 ns |
1.01 |
array/reductions/reduce/Float32/dims=1L |
617000 ns |
620375 ns |
0.99 |
array/reductions/reduce/Float32/dims=2 |
250792 ns |
251208 ns |
1.00 |
array/reductions/reduce/Float32/dims=2L |
477959 ns |
471792 ns |
1.01 |
array/reductions/reduce/Int64/1d |
457125 ns |
458333 ns |
1.00 |
array/reductions/reduce/Int64/dims=1 |
637584 ns |
641333 ns |
0.99 |
array/reductions/reduce/Int64/dims=1L |
1000042 ns |
1010167 ns |
0.99 |
array/reductions/reduce/Int64/dims=2 |
260084 ns |
257625 ns |
1.01 |
array/reductions/reduce/Int64/dims=2L |
664709 ns |
670834 ns |
0.99 |
array/shared/copy |
128416 ns |
130250 ns |
0.99 |
array/shared/copyto!/cpu_to_gpu |
37791 ns |
37291 ns |
1.01 |
array/shared/copyto!/gpu_to_cpu |
36750 ns |
37875 ns |
0.97 |
array/shared/copyto!/gpu_to_gpu |
37083 ns |
38209 ns |
0.97 |
array/shared/iteration/findall/bool |
1059042 ns |
1048458 ns |
1.01 |
array/shared/iteration/findall/int |
1218917 ns |
1223209 ns |
1.00 |
array/shared/iteration/findfirst/bool |
980667 ns |
976875 ns |
1.00 |
array/shared/iteration/findfirst/int |
991125 ns |
981250 ns |
1.01 |
array/shared/iteration/findmin/1d |
1049625 ns |
1050875 ns |
1.00 |
array/shared/iteration/findmin/2d |
1019834 ns |
1027416 ns |
0.99 |
array/shared/iteration/logical |
1522500 ns |
1525334 ns |
1.00 |
array/shared/iteration/scalar |
76.375 ns |
75.89014373716633 ns |
1.01 |
array/sorting/1d |
1913500 ns |
1921125 ns |
1.00 |
array/sorting/2d |
8265125 ns |
8319917 ns |
0.99 |
integration/byval/reference |
1116125 ns |
1120625 ns |
1.00 |
integration/byval/slices=1 |
1124542 ns |
1122291 ns |
1.00 |
integration/byval/slices=2 |
2031375 ns |
2029542 ns |
1.00 |
integration/byval/slices=3 |
6856667 ns |
6806458 ns |
1.01 |
integration/metaldevrt |
395125 ns |
396167 ns |
1.00 |
kernel/indexing |
204708 ns |
207625 ns |
0.99 |
kernel/indexing_checked |
389250 ns |
395208 ns |
0.98 |
kernel/launch |
2051 ns |
2101.8888888888887 ns |
0.98 |
kernel/rand |
406708 ns |
408958 ns |
0.99 |
latency/import |
1863992125 ns |
1862270958 ns |
1.00 |
latency/precompile |
33417766792 ns |
31631620958 ns |
1.06 |
latency/ttfp |
2293500584 ns |
2295776542 ns |
1.00 |
metal/synchronization/context |
529.3842105263158 ns |
525.958115183246 ns |
1.01 |
metal/synchronization/stream |
629.3604651162791 ns |
657.375796178344 ns |
0.96 |
This comment was automatically generated by workflow using github-action-benchmark.
GPUCompiler implements 8- and 16-bit atomics as operations on the containing 32-bit word, like LLVM's AtomicExpand, which assumes that the whole word can be accessed. That holds for device buffers, which Metal allocates in pages, so make it hold for threadgroup arrays too.
GPUCompiler implements 8- and 16-bit atomics on the containing 32-bit word. That word always exists in memory, as Metal allocates buffers in pages, but Metal's shader validation only allows accesses within the buffer's length, and drops the others: a UInt16 atomic on a 2-byte array had no effect. Pad the buffers that aren't backed by host memory to whole words, as threadgroup arrays are.
GPUCompiler now lowers LLVM atomics for Metal (JuliaGPU/GPUCompiler.jl#942), like an LLVM back-end would: selecting AIR's atomic intrinsics, expanding what AIR lacks (8- and 16-bit operations, compare-exchange loops for operations like nand or floating-point min/max), and implementing orderings with fences before MSL 4.1. Emit the atomic functions as LLVM atomics through UnsafeAtomics, with the device or workgroup scope for device or threadgroup memory, as MSL does. This supports ordered atomics on MSL 3.2 and 4.0, threadgroup Float32 add/sub on every version, and ordered 64-bit min/max. Orders passed as values are dispatched on, so that inference stays precise. Loads and stores only take the part of an order they can have, like Clang. Calls with explicit memory flags still use the MSL 4.1 intrinsics, whose ABI GPUCompiler now legalizes for older targets, so Metal.jl no longer does. This needs GPUCompiler 2.10 for the lowering; 2.11 is the first release that supports LLVM.jl 10.
GPUCompiler now sets it on every atomic load and read-modify-write it selects (without it, Apple's compiler turns read-modify-writes that don't change memory into loads that can be hoisted out of loops). Intrinsics called with explicit memory flags keep the volatile bit they are given.
MSL only makes device memory coherent within a threadgroup by default, so its relaxed loads are only guaranteed to observe stores from the same threadgroup; on an M1, a loop relaxed-loading a flag that another threadgroup sets never observes the store. LLVM's device-scope relaxed loads have to observe every thread's stores, which GPUCompiler now implements by making the ones that may be repeated acquire loads. Keep `atomic_load_explicit` with the relaxed order at MSL's semantics (and cost) by emitting it with the workgroup scope, and document how to wait for another threadgroup: with an acquire load, or a device-scope relaxed load through UnsafeAtomics. Test both, with bounded waits.
- The functions on LLVM atomics are generic over the element type and address space, instead of generated per combination. - `with_llvm_order` branches on the order once, for values and `Val`s. - The functions with memory flags come from one table, and call the Float32 intrinsics MSL uses instead of reinterpreting through UInt32. Like MSL, they then require Metal 4.1 for Float32 atomics on threadgroup memory (the functions without flags support those on every version). - One compare-exchange loop implements `atomic_fetch_op_explicit`, with and without memory flags.
GPUCompiler tests the AIR it selects for LLVM atomics, and how it legalizes the MSL 4.1 intrinsics, for every target. Only check the intrinsics Metal.jl calls itself, and share the list of targets.
MSL's atomic functions require the memory order to be a compile-time constant: its headers check this with METAL_CONST_ARG, and the AIR intrinsics take the order as an immediate. So did Metal.jl before this PR, by converting the order to a Val. Passing an order as a run-time value only worked because `with_llvm_order` branched on it. Require constants again. `llvm_order` maps a constant `memory_order` to an UnsafeAtomics ordering, and an order that isn't a constant is a dynamic call, i.e. an InvalidIRError. The default orders are Vals, so they don't depend on the functions being inlined, which GPUCompiler's interpreter doesn't always do.
On Julia 1.11, calling `fill` while generating the atomic functions resolves Metal's `fill` binding to Base's, so that defining it later in array.jl fails to precompile.
On Julia 1.11, GPUCompiler rejects the boxed constant that a run-time order needs with an ErrorException, before validating the IR.
GPUCompiler then stacks its shared method table below Metal's tables, whose overlays apply to code for every back-end.
GPUCompiler reads MSL's memory flags from the synchronization scope of an LLVM atomic or fence (e.g., device-mem-global), as LLVM's orderings order all memory. Emit the atomic functions and fences that take flags as LLVM atomics and fences too, so that they are lowered like the others, instead of calling the MSL 4.1 intrinsics. They then work on every target with ordered atomics (MSL 3.2), and threadgroup Float32 atomics take flags on every version. Relaxed fences, which LLVM cannot express, still call the intrinsic.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on JuliaGPU/GPUCompiler.jl#942, which lowers LLVM atomics for Metal (GPUCompiler 2.10), and requires GPUCompiler 2.12 and UnsafeAtomics 0.4 (unreleased):
singlethreadto the system scope, unless GPUCompiler compiles the code. It detects that with an overlay in the method table that Add a method table with overlays for every device back-end GPUCompiler.jl#978 adds for all back-ends, so Metal.jl declares its method tables withmethod_tablesinstead of building the view itself.Metal.jl's atomic functions now emit plain LLVM atomics through UnsafeAtomics, with an ordering and a synchronization scope (device scope for device memory, workgroup scope for threadgroup memory, as in MSL). GPUCompiler then does what an LLVM back-end would: it selects AIR's atomic intrinsics, expands what AIR lacks, and implements orderings with fences before MSL 4.1. That brings:
Float32add/sub on every MSL version;min/max.Other changes:
METAL_CONST_ARG, and whose AIR intrinsics take the order as an immediate). An order that isn't a constant fails to compile.device-mem-globalforMemoryFlagDevice, which GPUCompiler turns into the flags operand (Rename synchronization scopes for the target GPUCompiler.jl#979). Like the order, the flags have to be compile-time constants. Atomics with flags therefore work from MSL 3.2 instead of 4.1, includingFloat32atomics on threadgroup memory.atomic_thread_fencealso emits an LLVM fence with the flags in its scope, so acquire and release fences work from MSL 3.2 too (as sequentially-consistent fences before 4.1). Relaxed fences still callair.atomic.fence, as LLVM has no relaxed fences.air.atomic.*calls are left, and GPUCompiler legalizes the ABI of those that other code emits for older targets, so the downgrading code infinish_ir!is gone.atomic_fetch_op_explicit, with and without flags.New tests cover the LLVM atomics on MSL 3.2, 4.0 and 4.1: contention, partword neighbours, the compare-exchange expansions, acquire/release message passing, waiting on another threadgroup, and the kernel from JuliaGPU/GPUCompiler.jl#934. Checks of the AIR that GPUCompiler selects are left to GPUCompiler's tests; only the intrinsics Metal.jl calls itself are checked here.
Tested on an M1 with GPUCompiler 2.11.1, LLVM.jl 10.0.0 and UnsafeAtomics 0.3.3, before the last two commits:
device/intrinsics/atomics,device/intrinsics/synchronizationandkernelabstractions, on Julia 1.10 and 1.12: 2439 pass, 8 broken.examples/gtkfailed because its test worker died, which happens when it runs on its own too.--validate(Metal API and shader validation):device/intrinsics/atomics,synchronization,array,poolandkernelabstractions, 3087 pass, 8 broken.The last two commits (method tables, memory flags in scopes) were tested on the M1 on Julia 1.12, with GPUCompiler 2.12.0 and UnsafeAtomics 0.4:
device/intrinsics/atomics(229 pass) anddevice/intrinsics/synchronization(19 pass).