Repository navigation
Validate atomics for the PTX target - #996
Conversation
Atomics reach the PTX back-end from CUDA.jl's atomic functions, UnsafeAtomics and Atomix, Enzyme and Julia's intrinsics, so check them on the IR instead of in any one front-end. The NVPTX back-end doesn't reject everything the target can't run: before sm_60 it silently drops the system scope of read-modify-writes, it emits `.sys` 128-bit atomics that ptxas rejects before PTX ISA 8.4, and it aborts on unknown synchronization scopes and on sequentially-consistent 128-bit loads and stores. Other unsupported atomics (128-bit ones before sm_90, misaligned ones, atomics on local memory) fail in the back-end or in ptxas, without pointing at the Julia code that performs them. Platforms can also reject system-scope atomics the ISA supports, e.g. Pascal GPUs under Windows. Front-ends report that through the new `system_atomics` target field.
4370ffc to
849839a
Compare
Don't you need to use fences to fix them? I think this is what the CUDA headers do. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #996 +/- ##
==========================================
+ Coverage 88.62% 88.72% +0.09%
==========================================
Files 30 30
Lines 6443 6490 +47
==========================================
+ Hits 5710 5758 +48
+ Misses 733 732 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Yes, and the back-end does that now. JuliaPackaging/Yggdrasil#15013 backported llvm/llvm-project#222449 and the pre-sm_70 part of llvm/llvm-project#201468, so on sm_50/sm_60 NVPTX emits e.g. |
|
Fantastic! |
Atomics reach the PTX back-end from CUDA.jl's atomic functions, UnsafeAtomics and Atomix, Enzyme, and Julia's atomic intrinsics. CUDA.jl checked the capabilities they need with static assertions in its atomic functions, so atomics used without them weren't checked at all. This validates them on the IR instead, like
validate_iralready does for Metal, so that every front-end gets anInvalidIRErrorpointing at the Julia code.The NVPTX back-end (23.1.2+1) doesn't reject everything the target can't run:
.sys128-bit atomics with PTX ISA 8.3, which ptxas rejects ("requires PTX ISA .version 8.4"), also on LLVM main;Platforms can also lack system-scope atomics that the ISA has: Windows rejects modules that use them on Pascal GPUs (JuliaGPU/CUDA.jl#3187), and Tegra GPUs only have them from sm_72. GPUCompiler doesn't know the platform, so front-ends report it through a new
PTXCompilerTargetfield,system_atomics(defaulttrue, part of the target hash).What is deliberately accepted:
membars);clusterscope before sm_90, which NVPTX maps to the block scope, because a cluster is a single block there;The back-end workaround (sequentially-consistent 128-bit loads and stores) is kept apart from the target rules, so it can be revisited when updating the back-end. Ordered atomics before sm_70 aren't checked: the back-end lowers them to relaxed or volatile accesses bracketed by
membars, like libcu++ does (23.1.2+1, JuliaPackaging/Yggdrasil#15013; 23.1.2+0 dropped the fences of read-modify-writes and rejected ordered loads and stores).This makes released Atomix (which uses the system scope) and Enzyme's atomic accumulation fail on sm_5x, where they were silently downgraded to device scope. LLVM 24 rejects those too.
CUDA.jl's side is in JuliaGPU/CUDA.jl#3350.