Skip to content

JVM Panama vector kernels accept slab-backed (offset) FP32 operands - #1175

Merged
michalharakal merged 1 commit into
chore/apidump-bitnet-planesfrom
feature/1173-panama-offset-operands
Aug 26, 2026
Merged

michalharakal merged 1 commit into
chore/apidump-bitnet-planesfrom
feature/1173-panama-offset-operands

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1173 (parent #932). Stacked on #1174 — merge that first (tick "delete branch" and this one retargets to develop).

Completes the #1146 performance story: the JVM Panama vector paths guarded on data as? FloatArrayTensorData, so a slab-backed operand — what a ForwardScope produces since #1145/#1146 — silently fell back to the common scalar loops. Correct, but the JVM lost SIMD exactly where the scope machinery is in use.

  • Every vector entry point now reads dense-FP32 windows (array + base offset) matching both plain array data and StorageFloatTensorData: vectorFloatBinary across all four broadcast shapes, vectorFloatUnary (relu &c.), silu, the reduce-all sum (same left-to-right accumulation order — bit-identical), and the FP32 matmul FloatArray path, whose kernel SPI already took offsets natively. BLAS keeps an offset-0 guard (its JNI shim takes whole arrays); slab windows route through the kernel SPI.
  • JvmVectorKernels.binaryFloat/unaryFloat gain default-0 offset parameters — no other caller changes.
  • Conv paths keep their array-only guard (falling back to the common path for slab inputs) — conv operands are not part of the decode-loop story; noted here rather than silently.
  • SlabOperandVectorPathTest: every vector-eligible shape driven with both operands sliced from a slab at nonzero offsets (length 67 — not a multiple of any vector species) and required to be bit-identical to Ambient — the tight net for off-by-offset reads, which fail as plausible garbage rather than crashes.

Full pr-gate green on the stack.

🤖 Generated with Claude Code

…rands

The JVM vector paths guarded on 'data as? FloatArrayTensorData', so a
slab-backed operand — what a ForwardScope produces since #1145/#1146 —
silently fell back to the common scalar loops: correct, but the JVM
lost SIMD exactly where the scope machinery is in use.

The vector entry points now read dense-FP32 windows (array + base
offset) matching both plain array data and StorageFloatTensorData:
vectorFloatBinary in all four broadcast shapes, vectorFloatUnary
(relu &c.), silu, the reduce-all sum (same left-to-right accumulation
order, offset base), and the FP32 matmul FloatArray path — whose kernel
SPI already took offsets; only BLAS keeps its offset-0 guard, since the
JNI shim takes whole arrays. JvmVectorKernels.binaryFloat/unaryFloat
gain default-0 offset parameters.

SlabOperandVectorPathTest drives every vector-eligible shape with both
operands sliced from a slab at nonzero offsets (length 67 — no species
multiple) and requires bit-identical results vs Ambient: the tightest
guard against off-by-offset reads, which produce plausible garbage
rather than crashes.

Closes #1173.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 0f811dd into chore/apidump-bitnet-planes Aug 26, 2026
@michalharakal
michalharakal deleted the feature/1173-panama-offset-operands branch August 26, 2026 15:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant