Skip to content

0.40.1 packed ops.transpose double-permutes kernel-order bytes: block-order convention is not type-visible (regression from #969) #971

Description

@michalharakal

Summary

The 0.40.1 hotfix (#969, f72a267b, fixing #968) changed the byte-order contract of packed-quant ops.transpose in a patch release: it now assumes its input bytes are in canonical row-major block order and physically permutes them. Any packed tensor whose bytes are already input-block-major (kernel order) — the convention downstream converters were built around, and the convention the engine's own Q5_1TensorData/Q5_0TensorData kdoc still documents — is silently double-permuted into garbage by the new transpose.

Confirmed downstream regression (SKaiNET-transformers, pin bump 0.40.0 → 0.40.1, A/B-verified with only the pin differing):
LinearProjectionPreTransposedTest.preTransposed_q5_1_matches_classic_path_and_fp32_reference fails on 0.40.1, passes on 0.40.0 — plus 3 real-checkpoint gemma failures on the same classic pack() + ops.transpose path.

Mechanism

Two block orderings exist for a 2-D [out, in] packed weight (grid = outDim × blocksPerRow):

  • canonical / row-major: flat block index o * blocksPerRow + b — what GGUF stores, what Q4_0Quantizer emits, what toFloatArray() assumes.
  • kernel-native / input-block-major: b * outDim + o — what every heap matmul kernel reads ((blockIdx * outputDim + o) * bytesPerBlock).

Nothing in the type system distinguishes them: both populations legitimately share the same Q*BlockTensorData types. 0.40.0's transpose (shape-swap only) was correct for kernel-native bytes and wrong for canonical (#968). 0.40.1's transpose (physical permutation) is correct for canonical and wrong for kernel-native (this issue). Whichever order transpose() guesses, the other producer population silently corrupts. The block-grid permutation is not an involution for non-square grids, so the double application does not cancel.

Note the tier is irrelevant (the #969 hypothesis space was misleading here): the engine's own NativeLazyTransposeGroundTruthReproTest builds canonical bytes, so the fix looks right on every tier; downstream fixtures built kernel-native bytes, so it fails on every tier.

Contract contradictions left in-tree (0.40.1)

  1. Q5_1TensorData.kt:24-27 / Q5_0TensorData.kt:21-22 kdoc: "blocks are input-block-major … and the CPU-ops lazy transpose is a pure shape swap" — now false on both counts.
  2. Kernel SPI kdoc (Q5_0MatmulKernel.kt:29): "Matches Q5_0BlockTensorData.packedData" — true only post-transpose.
  3. Heap tier reads kernel-native while the MemSeg tier reads canonical (JvmQuantizedVectorKernels.kt:629,816) and the MemSeg lazy transpose is still shape-swap-only (DefaultCpuOpsJvm.kt:211-226) — same formats, opposite byte orders per carrier.
  4. Q4_K-over-MemorySegment has two contradictory implementations: JvmQuantizedVectorKernels.kt:667 (canonical, dead code) vs Q4KMemSegMatmulKernel + q4k_matmul.c:245 (kernel-native, live).
  5. Engine tests use two competing shape conventions: Q8_0MatmulDispatchTest.kt:68 ([in,out] + kernel-native, no transpose) vs PackedMatmulDispatchTest.kt:130 ([out,in] + canonical + transpose).
  6. The GGUF streaming loader emits ne-order [in,out] shapes un-reversed (StreamingGgufParametersLoader.kt:96) while transposePackedBlocks derives the grid assuming [out,in] — transposing a verbatim-loaded GGUF tensor computes the wrong permutation or require-fails.
  7. toFloatArray() / get() / matmulGeneric are hard-wired canonical (PackedBlockStorage.kt:52-61) — any post-transpose (kernel-native) tensor silently dequantizes to a block-permuted matrix.

Also: transpose(transpose(W)) ≠ W for non-square block grids, and Linear.onForward (Linear.kt:76-77) now pays an uncached O(bytes) copy per forward on packed weights.

Downstream mitigation (landing in SKaiNET-transformers)

BlockQuantPacking.pack() (classic path) stops eagerly relayouting — canonical bytes verbatim, matching what 0.40.1's transpose expects; the pre-transposed path keeps the load-time relayout + marker and never transposes. So do not revert 0.40.1 — reverting reintroduces #968 for canonical producers.

Proposed structural fix (0.41)

One convention, one owner, explicit in the type, loud on violation:

  1. Type-visible layout: follow the narrow-float precedent (NarrowFloatInputMajorTensorData) — kernel-order bytes get a distinct type (e.g. KernelPackedWeightData); plain Q*BlockTensorData means canonical row-major, always.
  2. Engine-owned prepacking: prepackForMatmul(weight): KernelPackedWeightData (absorbs downstream relayoutRowMajorToBlockMajor copies).
  3. Weight-transposing matmul primitive (matmulWT(x, W[out,in]), ggml-mul_mat-style) instead of ops.transpose on packed data — a true transpose of block-quantized data is unrepresentable without requantization; the current op is a layout conversion wearing a shape label that lies. ops.transpose on packed data becomes a loud error.
  4. GGUF loader dim-order fix (ne reversal), delete the dead canonical Q4_K MemSeg kernel, correct the stale kdocs above.
  5. Cross-repo contract tests + downstream canary CI before tagging releases; written policy that packedData byte semantics are public API — changes are breaking, never a hotfix.

Full analysis with the complete producer/consumer census: internal postmortem "Packed-Quant Layout Postmortem" (2026-08-12).

Refs: #968, #969, SKaiNET-transformers#307.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions