Skip to content

Validate atomics for the PTX target - #996

Merged
maleadt merged 1 commit into
mainfrom
tb/ptx-atomic-validation
Oct 7, 2026
Merged

maleadt merged 1 commit into
mainfrom
tb/ptx-atomic-validation

Conversation

@maleadt

@maleadt maleadt commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Atomics reach the PTX back-end from CUDA.jl's atomic functions, UnsafeAtomics and Atomix, Enzyme, and Julia's atomic intrinsics. CUDA.jl checked the capabilities they need with static assertions in its atomic functions, so atomics used without them weren't checked at all. This validates them on the IR instead, like validate_ir already does for Metal, so that every front-end gets an InvalidIRError pointing at the Julia code.

The NVPTX back-end (23.1.2+1) doesn't reject everything the target can't run:

  • before sm_60, it silently drops the system scope of read-modify-writes and compare-exchanges (LLVM 24 makes this an error, [NVPTX] Error on atomics with system scope pre-sm60 llvm/llvm-project#222458, but without the Julia frames);
  • it emits .sys 128-bit atomics with PTX ISA 8.3, which ptxas rejects ("requires PTX ISA .version 8.4"), also on LLVM main;
  • it aborts the process on unknown synchronization scopes and on sequentially-consistent 128-bit atomic loads and stores (also on LLVM main);
  • other unsupported atomics (128-bit ones before sm_90 or PTX ISA 8.3, misaligned ones, atomics on local, constant or parameter memory) fail in the back-end or in ptxas, without pointing at the Julia code.

Platforms can also lack system-scope atomics that the ISA has: Windows rejects modules that use them on Pascal GPUs (JuliaGPU/CUDA.jl#3187), and Tegra GPUs only have them from sm_72. GPUCompiler doesn't know the platform, so front-ends report it through a new PTXCompilerTarget field, system_atomics (default true, part of the target hash).

What is deliberately accepted:

  • loads, stores and fences at system scope before sm_60, and on platforms without system-scope atomics: they don't take a scope before sm_70 (they become volatile accesses and membars);
  • system scope on shared memory, which is only visible within a block or cluster;
  • cluster scope before sm_90, which NVPTX maps to the block scope, because a cluster is a single block there;
  • operations PTX has no instruction for, which NVPTX expands to compare-exchange loops.

The back-end workaround (sequentially-consistent 128-bit loads and stores) is kept apart from the target rules, so it can be revisited when updating the back-end. Ordered atomics before sm_70 aren't checked: the back-end lowers them to relaxed or volatile accesses bracketed by membars, like libcu++ does (23.1.2+1, JuliaPackaging/Yggdrasil#15013; 23.1.2+0 dropped the fences of read-modify-writes and rejected ordered loads and stores).

This makes released Atomix (which uses the system scope) and Enzyme's atomic accumulation fail on sm_5x, where they were silently downgraded to device scope. LLVM 24 rejects those too.

CUDA.jl's side is in JuliaGPU/CUDA.jl#3350.

@maleadt maleadt added enhancement New feature or request ptx Stuff about the NVIDIA PTX back-end. labels Oct 6, 2026
@maleadt maleadt added enhancement New feature or request ptx Stuff about the NVIDIA PTX back-end. labels Oct 6, 2026
Atomics reach the PTX back-end from CUDA.jl's atomic functions, UnsafeAtomics and Atomix,
Enzyme and Julia's intrinsics, so check them on the IR instead of in any one front-end. The
NVPTX back-end doesn't reject everything the target can't run: before sm_60 it silently
drops the system scope of read-modify-writes, it emits `.sys` 128-bit atomics that ptxas
rejects before PTX ISA 8.4, and it aborts on unknown synchronization scopes and on
sequentially-consistent 128-bit loads and stores. Other unsupported atomics (128-bit ones
before sm_90, misaligned ones, atomics on local memory) fail in the back-end or in ptxas,
without pointing at the Julia code that performs them.

Platforms can also reject system-scope atomics the ISA supports, e.g. Pascal GPUs under
Windows. Front-ends report that through the new `system_atomics` target field.
@maleadt
maleadt force-pushed the tb/ptx-atomic-validation branch from 4370ffc to 849839a Compare October 6, 2026 17:01
@vchuravy

vchuravy commented Oct 6, 2026

Copy link
Copy Markdown
Member

Ordered atomics before sm_70 aren't checked:

Don't you need to use fences to fix them? I think this is what the CUDA headers do.

@vchuravy

vchuravy commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

@codecov

codecov Bot commented Oct 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 88.72%. Comparing base (e6ff2e6) to head (849839a).
⚠️ Report is 5 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #996      +/-   ##
==========================================
+ Coverage   88.62%   88.72%   +0.09%     
==========================================
  Files          30       30              
  Lines        6443     6490      +47     
==========================================
+ Hits         5710     5758      +48     
+ Misses        733      732       -1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt

maleadt commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Don't you need to use fences to fix them? I think this is what the CUDA headers do.

Yes, and the back-end does that now. JuliaPackaging/Yggdrasil#15013 backported llvm/llvm-project#222449 and the pre-sm_70 part of llvm/llvm-project#201468, so on sm_50/sm_60 NVPTX emits e.g. membar.gl; atom.gpu.global.add; membar.gl for an acq_rel RMW, ld.volatile; membar.gl for an acquire load, membar.cta; st.volatile for a block-scope release store, and membar.sys on both sides of a seq_cst load. That's the same lowering as libcu++ and your JuliaGPU/CUDA.jl#1644, just done once in the back-end so that UnsafeAtomics/Atomix/Enzyme get it too.

@maleadt
maleadt merged commit 910b5ef into main Oct 7, 2026
36 checks passed
@maleadt
maleadt deleted the tb/ptx-atomic-validation branch October 7, 2026 06:36
@vchuravy

vchuravy commented Oct 7, 2026

Copy link
Copy Markdown
Member

Fantastic!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request ptx Stuff about the NVIDIA PTX back-end.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants