Repository navigation
Conversation
LLVM implements 8- and 16-bit atomics as a compare-and-swap on the containing 32-bit word, which can extend past the end of an array. That never faults, since allocations are at least 256-byte aligned, but compute-sanitizer reports it as an out-of-bounds access and aborts the kernel. This affects any sub-word atomic going through LLVM: Int8 and Int16 operations from UnsafeAtomics, Atomix or KernelAbstractions on every architecture, and 16-bit operations before sm_70. Round pool allocations, static shared memory arrays and the dynamic shared memory size up to a multiple of 4 bytes, like Metal.jl does. The pool already hands out much larger blocks, so this costs no memory. Memory CUDA.jl did not allocate, such as pointers passed to unsafe_wrap, can't be padded; the docs now say so.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PTX has few 8- and 16-bit atomics, so LLVM implements them as a compare-and-swap on the containing 32-bit word. When the value sits in the last partial word of an array, that word extends past the end of the allocation. This never faults (allocations are 256-byte aligned), but compute-sanitizer reports it as an out-of-bounds access and aborts the kernel, leaving a sticky error:
This hits any sub-word atomic that goes through LLVM: Int8/Int16 read-modify-writes and CAS from UnsafeAtomics, Atomix or KernelAbstractions on every architecture (current LLVM never emits
atom.cas.b16), and CUDA.jl's own 16-bit atomics before sm_70 or BFloat16 add before sm_90. It applies to global memory and to static and dynamic shared memory alike. NVCC doesn't run into this because it emitsatom.cas.b16, which ptxas widens after the sanitizer's view of the access; libcu++'scuda::atomic_ref<short>has the same problem we do.Like Metal.jl did in JuliaGPU/Metal.jl#977, this pads pool allocations, static shared memory arrays and the dynamic shared memory size of a launch to a whole number of 32-bit words. That costs no memory: the stream-ordered pool hands out 512-byte blocks even for 2-byte requests, and shared memory is allocated in units of at least 128 bytes. The new tests make the sanitizer job abort without this change.
Caveats:
unsafe_wrap, the low-levelCUDA.alloc, device-sidemalloc). The kernel programming docs now say so.CUDA.@allocated CuArray{Int8}(undef, 1)reports 4 bytes.FUNC_ATTRIBUTE_MAX_DYNAMIC_SHARED_SIZE_BYTESlimit will now exceed it by up to 3 bytes.With this in place, #3350 shouldn't need its native
atom.cas.b16paths to keep the sanitizer happy, and 8-bit atomics (which #3350 enables through UnsafeAtomics and which have no native fallback) are covered.