Skip to content

Use the generic implementation for CUDA arrays - #87

Merged
maleadt merged 1 commit into
tb/device-scopefrom
tb/drop-cudacore-ext
Oct 6, 2026
Merged

maleadt merged 1 commit into
tb/device-scopefrom
tb/drop-cudacore-ext

Conversation

@maleadt

@maleadt maleadt commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

With UnsafeAtomics 0.4 (whose atomics GPUCompiler legalizes) and the device scope (previous PR), the generic implementation emits the same atomics for CUDA arrays as CUDA.jl's atomic_*! functions, and also honours the requested ordering, which the extension ignored. The same as #85 did for Metal. For @atomic A[i] += 1f0 it emits fence.sc.gpu; atom.acquire.gpu.global.add.f32, i.e. a sequentially consistent device-scope atomic.

CUDACore stays a weak dependency, with compat 6.5, so environments with an older CUDA.jl keep resolving Atomix 1.5 with the extension:

Until CUDA.jl 6.5 is out, the CUDA CI job can't resolve CUDACore, and nobody can install this version together with CUDACore. With the bound relaxed, the CUDA tests pass on the generic implementation with CUDACore 6.4.2 on Julia 1.10 and 1.12 (RTX 5080), as does KernelAbstractions' histogram example.

With UnsafeAtomics 0.4 and the device scope, the generic implementation emits
the atomics CUDA.jl would, and also honours the ordering, which the extension
ignored. CUDACore stays a weak dependency to keep older CUDA.jl versions on the
extension: before 6.2 they compile with Julia's own LLVM, which can't lower all
of these atomics, and before 6.5 their NVPTX back-end lowers float subtraction
to a compare-and-swap loop and doesn't order atomics on Pascal GPUs.
@maleadt
maleadt added this pull request to stack #88 October 6, 2026 09:27
@maleadt
maleadt merged commit 5bfda0b into main Oct 6, 2026
6 of 7 checks passed
@maleadt
maleadt deleted the tb/drop-cudacore-ext branch October 6, 2026 09:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant