fix(cpu-jvm): scoped dense-FP32 activations must not fall out of the quantized matmul chooser - #1211
Merged
Merged
Conversation
…quantized matmul chooser chooseQuantizedMatmul2D accepted only FloatArrayTensorData or MemorySegmentBackedData activations. Any other dense FP32 TensorData — most importantly the slab-backed StorageFloatTensorData a ScopedExecutionContext forward hands every op mid-step (#1145/#1146) — fell through 'else -> return null', and for a Q8 MemorySegment weight the eventual fallback is matmulGeneric, whose per-element get() on that data type returns the RAW quantization byte, not the value: silently wrong logits, the #993 class of bug, surfaced end-to-end when a downstream decode loop adopted forwardScope. Accept any dense activation (encoding == null) by copying it out through its own offset-aware copyToFloatArray; packed/encoded activations still bail to the adaptive path. Regression test pins a nonzero-offset slab activation x Q8 MemorySegment weight to the dense-activation result bit-for-bit (verified failing without the fix).
This was referenced Aug 29, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
chooseQuantizedMatmul2Daccepted onlyFloatArrayTensorDataorMemorySegmentBackedDataactivations. Any other dense FP32TensorData— most importantly the slab-backedStorageFloatTensorDataaScopedExecutionContextforward hands every op mid-step (#1145/#1146) — fell throughelse -> return null. For aQ8MemorySegmentTensorDataweight the eventual fallback ismatmulGeneric, whose per-elementget()on that data type returns the raw quantization byte, not the value: silently wrong logits (the #993 class of bug).Surfaced end-to-end in SKaiNET-transformers when the decode loop adopted
forwardScope(#343 there): a Q8-MemSeg-projected Qwen diverged by ~10 logits at step 0 while every dense-FP32 model stayed bit-identical.Fix
Accept any dense activation (
encoding == null) by copying it out through its own offset-awarecopyToFloatArray(); packed/encoded activations still bail to the adaptive path.Test
Q8 matmul with a slab-backed scoped activation matches the dense-activation result— a nonzero-offsetForwardScopeslab activation × Q8 MemSeg weight pinned bit-for-bit to the dense-activation result. Verified failing (3 assertion sites) with the fix reverted.🤖 Generated with Claude Code