JVM Panama vector kernels accept slab-backed (offset) FP32 operands - #1175
Merged
michalharakal merged 1 commit intoAug 26, 2026
Merged
michalharakal merged 1 commit into
michalharakal merged 1 commit into
Conversation
…rands The JVM vector paths guarded on 'data as? FloatArrayTensorData', so a slab-backed operand — what a ForwardScope produces since #1145/#1146 — silently fell back to the common scalar loops: correct, but the JVM lost SIMD exactly where the scope machinery is in use. The vector entry points now read dense-FP32 windows (array + base offset) matching both plain array data and StorageFloatTensorData: vectorFloatBinary in all four broadcast shapes, vectorFloatUnary (relu &c.), silu, the reduce-all sum (same left-to-right accumulation order, offset base), and the FP32 matmul FloatArray path — whose kernel SPI already took offsets; only BLAS keeps its offset-0 guard, since the JNI shim takes whole arrays. JvmVectorKernels.binaryFloat/unaryFloat gain default-0 offset parameters. SlabOperandVectorPathTest drives every vector-eligible shape with both operands sliced from a slab at nonzero offsets (length 67 — no species multiple) and requires bit-identical results vs Ambient: the tightest guard against off-by-offset reads, which produce plausible garbage rather than crashes. Closes #1173. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1173 (parent #932). Stacked on #1174 — merge that first (tick "delete branch" and this one retargets to
develop).Completes the #1146 performance story: the JVM Panama vector paths guarded on
data as? FloatArrayTensorData, so a slab-backed operand — what aForwardScopeproduces since #1145/#1146 — silently fell back to the common scalar loops. Correct, but the JVM lost SIMD exactly where the scope machinery is in use.StorageFloatTensorData:vectorFloatBinaryacross all four broadcast shapes,vectorFloatUnary(relu&c.),silu, the reduce-all sum (same left-to-right accumulation order — bit-identical), and the FP32 matmul FloatArray path, whose kernel SPI already took offsets natively. BLAS keeps an offset-0 guard (its JNI shim takes whole arrays); slab windows route through the kernel SPI.JvmVectorKernels.binaryFloat/unaryFloatgain default-0 offset parameters — no other caller changes.SlabOperandVectorPathTest: every vector-eligible shape driven with both operands sliced from a slab at nonzero offsets (length 67 — not a multiple of any vector species) and required to be bit-identical to Ambient — the tight net for off-by-offset reads, which fail as plausible garbage rather than crashes.Full pr-gate green on the stack.
🤖 Generated with Claude Code