fix: Q6_K/Q5_1/Q5_0 matmul silently falls through for MemorySegment-backed activations - #992
Merged
Merged
Conversation
…acked activations chooseQuantizedMatmulHeap (the fallback dispatcher DefaultCpuOpsJvm. chooseQuantizedMatmul intentionally defers to for Q6_K/Q5_1/Q5_0) required the activation tensor to be exactly FloatArrayTensorData, silently returning null otherwise. Real forward-pass activations produced with DirectCpuExecutionContext(tensorDataFactory = MemorySegmentTensorDataFactory()) - the configuration every production caller uses - are MemorySegmentTensorData, not FloatArrayTensorData, so they never matched. Dispatch fell through the unguarded matmulGeneric fallback, which has no packed-quant handling and threw ClassCastException: Byte cannot be cast to Float reading raw packed bytes. Use the universal TensorData.copyToFloatArray() instead of a type-gated cast; FloatArrayTensorData already overrides it with a cheap buffer.copyOf(), so the fast path is unaffected. Also flattens leading batch/sequence dimensions in DefaultCpuOpsJvm.chooseQuantizedMatmul before its own rank-2 fast-path check, since linearProject's `[..., in]` inputs are a real (if not the triggering) case that fast path can't otherwise reach. Fixes #991.
aharakal
previously approved these changes
Aug 14, 2026
…o a broken fallback chooseQuantizedMatmulHeap (DefaultCpuOps.kt) and chooseQuantizedMatmul (DefaultCpuOpsJvm.kt) both required a.shape.rank >= 2, so a rank-1 activation — the single-token hidden-state vector a real forward pass produces once incremental decode moves past the initial batched prefill — skipped the packed-quant kernel dispatch entirely and fell through to matmulGeneric's untyped per-element TensorData.get(). For a packed-quant (or PreTransposedWeight-wrapped, e.g. PreTransposedQ4_K) weight, that get() returns the raw packed byte rather than a dequantized Float, throwing "ClassCastException: Byte cannot be cast to Float" on every decode step past the first token — reproduced end-to-end against a real Llama-3.2-1B GGUF via SKaiNET-transformers' KLlamaJava. Both guards now admit rank-1 (`a.shape.rank < 1`, down from `< 2`); the existing leading-dim-flattening logic already generalizes correctly to it (empty leading dims, flatBatch=1) without further changes. Also hardens matmulGeneric itself as defense in depth: a lazily materialized copyToFloatArray() fallback if a TensorData's generic get() ever returns something other than the tensor's own dtype, instead of an unconditional unsafe `as Float` cast. Adds a regression test exercising a rank-1 FP32 activation against a Q4_K-packed weight end to end. Co-authored-by: Claude <noreply@anthropic.com>
|
Pushed an additional fix on top of this one, filed as #993 — a related but distinct dispatch gap in the same functions this PR touches ( Verified end-to-end against a real Llama-3.2-1B GGUF via SKaiNET-transformers' |
aharakal
approved these changes
Aug 16, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
chooseQuantizedMatmulHeap(the fallback dispatcherDefaultCpuOpsJvm.chooseQuantizedMatmulintentionally defers to for Q6_K/Q5_1/Q5_0 — see the "handled in DefaultCpuOpsBase via the kernel registry... not intercepted here" comment) required the activation tensor to be exactlyFloatArrayTensorData, silently returningnullfor anything else.Real forward-pass activations produced with
DirectCpuExecutionContext(tensorDataFactory = MemorySegmentTensorDataFactory())— the configuration every production caller uses, includingKLlamaJava.loadGGUFin SKaiNET-transformers — areMemorySegmentTensorData, notFloatArrayTensorData, so they never matched this check. Dispatch fell through to the unguardedmatmulGenericfallback, which has no packed-quant handling at all and threw:reading raw packed Q6_K bytes as if they were
Float.Fix
chooseQuantizedMatmulHeap2Dnow calls the universalTensorData.copyToFloatArray()instead of gating on theFloatArrayTensorDatasubtype.FloatArrayTensorDataalready overrides that method with a cheapbuffer.copyOf(), so the existing fast path is unaffected — this only adds correct handling for everything else (MemorySegment-backed, etc.).DefaultCpuOpsJvm.chooseQuantizedMatmulnow flattens leading batch/sequence dimensions before its rank-2 fast-path check. Not the trigger for this specific bug (the repro's activation tensor was already rank-2), butlinearProject's documented[..., in]input shape is a real case this fast path couldn't otherwise reach, so it's included as a related robustness fix.QuantizedMemSegMatmulTest) that reproduces the exact scenario: a Q6_K weight matmul'd against aMemorySegmentTensorDataFactory-backed FP32 activation, verifying it returns a correctly-shaped, finite result instead of throwing.Verification
Beyond the new unit test, I reproduced and verified the fix against a real model end-to-end (Llama-3.2-1B-Instruct-Q4_K_M.gguf via
KLlamaJava.loadGGUF, from a downstream KMP app integrating SKaiNET-transformers 0.40.2). Before the fix: crash on the first forward pass. After: coherent generated text, in ~5s total including model load, via theNATIVE_OPTIMIZEDpacked-quant path (not a slow FP32-dequant fallback).Fixes #991.