Skip to content

fix: Q6_K/Q5_1/Q5_0 matmul silently falls through for MemorySegment-backed activations - #992

Merged
michalharakal merged 2 commits into
developfrom
fix/q6k-memseg-matmul-dispatch
Aug 16, 2026
Merged

michalharakal merged 2 commits into
developfrom
fix/q6k-memseg-matmul-dispatch

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

chooseQuantizedMatmulHeap (the fallback dispatcher DefaultCpuOpsJvm.chooseQuantizedMatmul intentionally defers to for Q6_K/Q5_1/Q5_0 — see the "handled in DefaultCpuOpsBase via the kernel registry... not intercepted here" comment) required the activation tensor to be exactly FloatArrayTensorData, silently returning null for anything else.

Real forward-pass activations produced with DirectCpuExecutionContext(tensorDataFactory = MemorySegmentTensorDataFactory()) — the configuration every production caller uses, including KLlamaJava.loadGGUF in SKaiNET-transformers — are MemorySegmentTensorData, not FloatArrayTensorData, so they never matched this check. Dispatch fell through to the unguarded matmulGeneric fallback, which has no packed-quant handling at all and threw:

java.lang.ClassCastException: class java.lang.Byte cannot be cast to class java.lang.Float
	at sk.ainet.exec.tensor.ops.DefaultCpuOpsBase.matmulGeneric$lambda$1(DefaultCpuOps.kt:726)

reading raw packed Q6_K bytes as if they were Float.

Fix

  • chooseQuantizedMatmulHeap2D now calls the universal TensorData.copyToFloatArray() instead of gating on the FloatArrayTensorData subtype. FloatArrayTensorData already overrides that method with a cheap buffer.copyOf(), so the existing fast path is unaffected — this only adds correct handling for everything else (MemorySegment-backed, etc.).
  • DefaultCpuOpsJvm.chooseQuantizedMatmul now flattens leading batch/sequence dimensions before its rank-2 fast-path check. Not the trigger for this specific bug (the repro's activation tensor was already rank-2), but linearProject's documented [..., in] input shape is a real case this fast path couldn't otherwise reach, so it's included as a related robustness fix.
  • Added a regression test (QuantizedMemSegMatmulTest) that reproduces the exact scenario: a Q6_K weight matmul'd against a MemorySegmentTensorDataFactory-backed FP32 activation, verifying it returns a correctly-shaped, finite result instead of throwing.

Verification

Beyond the new unit test, I reproduced and verified the fix against a real model end-to-end (Llama-3.2-1B-Instruct-Q4_K_M.gguf via KLlamaJava.loadGGUF, from a downstream KMP app integrating SKaiNET-transformers 0.40.2). Before the fix: crash on the first forward pass. After: coherent generated text, in ~5s total including model load, via the NATIVE_OPTIMIZED packed-quant path (not a slow FP32-dequant fallback).

Fixes #991.

…acked activations

chooseQuantizedMatmulHeap (the fallback dispatcher DefaultCpuOpsJvm.
chooseQuantizedMatmul intentionally defers to for Q6_K/Q5_1/Q5_0) required
the activation tensor to be exactly FloatArrayTensorData, silently returning
null otherwise. Real forward-pass activations produced with
DirectCpuExecutionContext(tensorDataFactory = MemorySegmentTensorDataFactory())
- the configuration every production caller uses - are MemorySegmentTensorData,
not FloatArrayTensorData, so they never matched. Dispatch fell through the
unguarded matmulGeneric fallback, which has no packed-quant handling and threw
ClassCastException: Byte cannot be cast to Float reading raw packed bytes.

Use the universal TensorData.copyToFloatArray() instead of a type-gated cast;
FloatArrayTensorData already overrides it with a cheap buffer.copyOf(), so the
fast path is unaffected. Also flattens leading batch/sequence dimensions in
DefaultCpuOpsJvm.chooseQuantizedMatmul before its own rank-2 fast-path check,
since linearProject's `[..., in]` inputs are a real (if not the triggering)
case that fast path can't otherwise reach.

Fixes #991.
aharakal
aharakal previously approved these changes Aug 14, 2026
…o a broken fallback

chooseQuantizedMatmulHeap (DefaultCpuOps.kt) and chooseQuantizedMatmul
(DefaultCpuOpsJvm.kt) both required a.shape.rank >= 2, so a rank-1
activation — the single-token hidden-state vector a real forward pass
produces once incremental decode moves past the initial batched prefill —
skipped the packed-quant kernel dispatch entirely and fell through to
matmulGeneric's untyped per-element TensorData.get(). For a packed-quant
(or PreTransposedWeight-wrapped, e.g. PreTransposedQ4_K) weight, that
get() returns the raw packed byte rather than a dequantized Float,
throwing "ClassCastException: Byte cannot be cast to Float" on every
decode step past the first token — reproduced end-to-end against a real
Llama-3.2-1B GGUF via SKaiNET-transformers' KLlamaJava.

Both guards now admit rank-1 (`a.shape.rank < 1`, down from `< 2`); the
existing leading-dim-flattening logic already generalizes correctly to
it (empty leading dims, flatBatch=1) without further changes.

Also hardens matmulGeneric itself as defense in depth: a lazily
materialized copyToFloatArray() fallback if a TensorData's generic
get() ever returns something other than the tensor's own dtype, instead
of an unconditional unsafe `as Float` cast.

Adds a regression test exercising a rank-1 FP32 activation against a
Q4_K-packed weight end to end.

Co-authored-by: Claude <noreply@anthropic.com>
@fiwio-developer

Copy link
Copy Markdown

Pushed an additional fix on top of this one, filed as #993 — a related but distinct dispatch gap in the same functions this PR touches (chooseQuantizedMatmulHeap / chooseQuantizedMatmul): both required a.shape.rank >= 2, so a rank-1 activation (the single-token hidden-state vector every real forward pass produces once incremental decode moves past the initial batched prefill) skipped the packed-quant kernel dispatch entirely and hit the same matmulGeneric ClassCastException this PR already fixes for the rank-2 MemorySegment case.

Verified end-to-end against a real Llama-3.2-1B GGUF via SKaiNET-transformers' KLlamaJava — this PR's original fix alone was not sufficient to get past the first post-prefill decode step; both fixes together are.

@michalharakal
michalharakal requested a review from aharakal August 16, 2026 10:52
@michalharakal
michalharakal merged commit c130f89 into develop Aug 16, 2026
13 checks passed
@michalharakal
michalharakal deleted the fix/q6k-memseg-matmul-dispatch branch August 16, 2026 18:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants