Skip to content

fix(cpu-jvm): scoped dense-FP32 activations must not fall out of the quantized matmul chooser - #1211

Merged
michalharakal merged 1 commit into
developfrom
fix/scoped-activation-quant-matmul
Aug 29, 2026
Merged

michalharakal merged 1 commit into
developfrom
fix/scoped-activation-quant-matmul

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Problem

chooseQuantizedMatmul2D accepted only FloatArrayTensorData or MemorySegmentBackedData activations. Any other dense FP32 TensorData — most importantly the slab-backed StorageFloatTensorData a ScopedExecutionContext forward hands every op mid-step (#1145/#1146) — fell through else -> return null. For a Q8MemorySegmentTensorData weight the eventual fallback is matmulGeneric, whose per-element get() on that data type returns the raw quantization byte, not the value: silently wrong logits (the #993 class of bug).

Surfaced end-to-end in SKaiNET-transformers when the decode loop adopted forwardScope (#343 there): a Q8-MemSeg-projected Qwen diverged by ~10 logits at step 0 while every dense-FP32 model stayed bit-identical.

Fix

Accept any dense activation (encoding == null) by copying it out through its own offset-aware copyToFloatArray(); packed/encoded activations still bail to the adaptive path.

Test

Q8 matmul with a slab-backed scoped activation matches the dense-activation result — a nonzero-offset ForwardScope slab activation × Q8 MemSeg weight pinned bit-for-bit to the dense-activation result. Verified failing (3 assertion sites) with the fix reverted.

🤖 Generated with Claude Code

…quantized matmul chooser

chooseQuantizedMatmul2D accepted only FloatArrayTensorData or
MemorySegmentBackedData activations. Any other dense FP32 TensorData —
most importantly the slab-backed StorageFloatTensorData a
ScopedExecutionContext forward hands every op mid-step (#1145/#1146) —
fell through 'else -> return null', and for a Q8 MemorySegment weight
the eventual fallback is matmulGeneric, whose per-element get() on that
data type returns the RAW quantization byte, not the value: silently
wrong logits, the #993 class of bug, surfaced end-to-end when a
downstream decode loop adopted forwardScope.

Accept any dense activation (encoding == null) by copying it out
through its own offset-aware copyToFloatArray; packed/encoded
activations still bail to the adaptive path. Regression test pins a
nonzero-offset slab activation x Q8 MemorySegment weight to the
dense-activation result bit-for-bit (verified failing without the fix).
@michalharakal
michalharakal merged commit 02abbed into develop Aug 29, 2026
16 checks passed
@michalharakal
michalharakal deleted the fix/scoped-activation-quant-matmul branch August 29, 2026 18:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant