Summary
The 0.40.1 hotfix (#969, f72a267b, fixing #968) changed the byte-order contract of packed-quant ops.transpose in a patch release: it now assumes its input bytes are in canonical row-major block order and physically permutes them. Any packed tensor whose bytes are already input-block-major (kernel order) — the convention downstream converters were built around, and the convention the engine's own Q5_1TensorData/Q5_0TensorData kdoc still documents — is silently double-permuted into garbage by the new transpose.
Confirmed downstream regression (SKaiNET-transformers, pin bump 0.40.0 → 0.40.1, A/B-verified with only the pin differing):
LinearProjectionPreTransposedTest.preTransposed_q5_1_matches_classic_path_and_fp32_reference fails on 0.40.1, passes on 0.40.0 — plus 3 real-checkpoint gemma failures on the same classic pack() + ops.transpose path.
Mechanism
Two block orderings exist for a 2-D [out, in] packed weight (grid = outDim × blocksPerRow):
- canonical / row-major: flat block index
o * blocksPerRow + b — what GGUF stores, what Q4_0Quantizer emits, what toFloatArray() assumes.
- kernel-native / input-block-major:
b * outDim + o — what every heap matmul kernel reads ((blockIdx * outputDim + o) * bytesPerBlock).
Nothing in the type system distinguishes them: both populations legitimately share the same Q*BlockTensorData types. 0.40.0's transpose (shape-swap only) was correct for kernel-native bytes and wrong for canonical (#968). 0.40.1's transpose (physical permutation) is correct for canonical and wrong for kernel-native (this issue). Whichever order transpose() guesses, the other producer population silently corrupts. The block-grid permutation is not an involution for non-square grids, so the double application does not cancel.
Note the tier is irrelevant (the #969 hypothesis space was misleading here): the engine's own NativeLazyTransposeGroundTruthReproTest builds canonical bytes, so the fix looks right on every tier; downstream fixtures built kernel-native bytes, so it fails on every tier.
Contract contradictions left in-tree (0.40.1)
Q5_1TensorData.kt:24-27 / Q5_0TensorData.kt:21-22 kdoc: "blocks are input-block-major … and the CPU-ops lazy transpose is a pure shape swap" — now false on both counts.
- Kernel SPI kdoc (
Q5_0MatmulKernel.kt:29): "Matches Q5_0BlockTensorData.packedData" — true only post-transpose.
- Heap tier reads kernel-native while the MemSeg tier reads canonical (
JvmQuantizedVectorKernels.kt:629,816) and the MemSeg lazy transpose is still shape-swap-only (DefaultCpuOpsJvm.kt:211-226) — same formats, opposite byte orders per carrier.
- Q4_K-over-MemorySegment has two contradictory implementations:
JvmQuantizedVectorKernels.kt:667 (canonical, dead code) vs Q4KMemSegMatmulKernel + q4k_matmul.c:245 (kernel-native, live).
- Engine tests use two competing shape conventions:
Q8_0MatmulDispatchTest.kt:68 ([in,out] + kernel-native, no transpose) vs PackedMatmulDispatchTest.kt:130 ([out,in] + canonical + transpose).
- The GGUF streaming loader emits ne-order
[in,out] shapes un-reversed (StreamingGgufParametersLoader.kt:96) while transposePackedBlocks derives the grid assuming [out,in] — transposing a verbatim-loaded GGUF tensor computes the wrong permutation or require-fails.
toFloatArray() / get() / matmulGeneric are hard-wired canonical (PackedBlockStorage.kt:52-61) — any post-transpose (kernel-native) tensor silently dequantizes to a block-permuted matrix.
Also: transpose(transpose(W)) ≠ W for non-square block grids, and Linear.onForward (Linear.kt:76-77) now pays an uncached O(bytes) copy per forward on packed weights.
Downstream mitigation (landing in SKaiNET-transformers)
BlockQuantPacking.pack() (classic path) stops eagerly relayouting — canonical bytes verbatim, matching what 0.40.1's transpose expects; the pre-transposed path keeps the load-time relayout + marker and never transposes. So do not revert 0.40.1 — reverting reintroduces #968 for canonical producers.
Proposed structural fix (0.41)
One convention, one owner, explicit in the type, loud on violation:
- Type-visible layout: follow the narrow-float precedent (
NarrowFloatInputMajorTensorData) — kernel-order bytes get a distinct type (e.g. KernelPackedWeightData); plain Q*BlockTensorData means canonical row-major, always.
- Engine-owned prepacking:
prepackForMatmul(weight): KernelPackedWeightData (absorbs downstream relayoutRowMajorToBlockMajor copies).
- Weight-transposing matmul primitive (
matmulWT(x, W[out,in]), ggml-mul_mat-style) instead of ops.transpose on packed data — a true transpose of block-quantized data is unrepresentable without requantization; the current op is a layout conversion wearing a shape label that lies. ops.transpose on packed data becomes a loud error.
- GGUF loader dim-order fix (ne reversal), delete the dead canonical Q4_K MemSeg kernel, correct the stale kdocs above.
- Cross-repo contract tests + downstream canary CI before tagging releases; written policy that
packedData byte semantics are public API — changes are breaking, never a hotfix.
Full analysis with the complete producer/consumer census: internal postmortem "Packed-Quant Layout Postmortem" (2026-08-12).
Refs: #968, #969, SKaiNET-transformers#307.
Summary
The 0.40.1 hotfix (#969,
f72a267b, fixing #968) changed the byte-order contract of packed-quantops.transposein a patch release: it now assumes its input bytes are in canonical row-major block order and physically permutes them. Any packed tensor whose bytes are already input-block-major (kernel order) — the convention downstream converters were built around, and the convention the engine's ownQ5_1TensorData/Q5_0TensorDatakdoc still documents — is silently double-permuted into garbage by the new transpose.Confirmed downstream regression (SKaiNET-transformers, pin bump 0.40.0 → 0.40.1, A/B-verified with only the pin differing):
LinearProjectionPreTransposedTest.preTransposed_q5_1_matches_classic_path_and_fp32_referencefails on 0.40.1, passes on 0.40.0 — plus 3 real-checkpoint gemma failures on the same classicpack()+ops.transposepath.Mechanism
Two block orderings exist for a 2-D
[out, in]packed weight (grid =outDim × blocksPerRow):o * blocksPerRow + b— what GGUF stores, whatQ4_0Quantizeremits, whattoFloatArray()assumes.b * outDim + o— what every heap matmul kernel reads ((blockIdx * outputDim + o) * bytesPerBlock).Nothing in the type system distinguishes them: both populations legitimately share the same
Q*BlockTensorDatatypes. 0.40.0's transpose (shape-swap only) was correct for kernel-native bytes and wrong for canonical (#968). 0.40.1's transpose (physical permutation) is correct for canonical and wrong for kernel-native (this issue). Whichever ordertranspose()guesses, the other producer population silently corrupts. The block-grid permutation is not an involution for non-square grids, so the double application does not cancel.Note the tier is irrelevant (the #969 hypothesis space was misleading here): the engine's own
NativeLazyTransposeGroundTruthReproTestbuilds canonical bytes, so the fix looks right on every tier; downstream fixtures built kernel-native bytes, so it fails on every tier.Contract contradictions left in-tree (0.40.1)
Q5_1TensorData.kt:24-27/Q5_0TensorData.kt:21-22kdoc: "blocks are input-block-major … and the CPU-ops lazy transpose is a pure shape swap" — now false on both counts.Q5_0MatmulKernel.kt:29): "Matches Q5_0BlockTensorData.packedData" — true only post-transpose.JvmQuantizedVectorKernels.kt:629,816) and the MemSeg lazy transpose is still shape-swap-only (DefaultCpuOpsJvm.kt:211-226) — same formats, opposite byte orders per carrier.JvmQuantizedVectorKernels.kt:667(canonical, dead code) vsQ4KMemSegMatmulKernel+q4k_matmul.c:245(kernel-native, live).Q8_0MatmulDispatchTest.kt:68([in,out]+ kernel-native, no transpose) vsPackedMatmulDispatchTest.kt:130([out,in]+ canonical + transpose).[in,out]shapes un-reversed (StreamingGgufParametersLoader.kt:96) whiletransposePackedBlocksderives the grid assuming[out,in]— transposing a verbatim-loaded GGUF tensor computes the wrong permutation orrequire-fails.toFloatArray()/get()/matmulGenericare hard-wired canonical (PackedBlockStorage.kt:52-61) — any post-transpose (kernel-native) tensor silently dequantizes to a block-permuted matrix.Also:
transpose(transpose(W)) ≠ Wfor non-square block grids, andLinear.onForward(Linear.kt:76-77) now pays an uncached O(bytes) copy per forward on packed weights.Downstream mitigation (landing in SKaiNET-transformers)
BlockQuantPacking.pack()(classic path) stops eagerly relayouting — canonical bytes verbatim, matching what 0.40.1's transpose expects; the pre-transposed path keeps the load-time relayout + marker and never transposes. So do not revert 0.40.1 — reverting reintroduces #968 for canonical producers.Proposed structural fix (0.41)
One convention, one owner, explicit in the type, loud on violation:
NarrowFloatInputMajorTensorData) — kernel-order bytes get a distinct type (e.g.KernelPackedWeightData); plainQ*BlockTensorDatameans canonical row-major, always.prepackForMatmul(weight): KernelPackedWeightData(absorbs downstreamrelayoutRowMajorToBlockMajorcopies).matmulWT(x, W[out,in]), ggml-mul_mat-style) instead ofops.transposeon packed data — a true transpose of block-quantized data is unrepresentable without requantization; the current op is a layout conversion wearing a shape label that lies.ops.transposeon packed data becomes a loud error.packedDatabyte semantics are public API — changes are breaking, never a hotfix.Full analysis with the complete producer/consumer census: internal postmortem "Packed-Quant Layout Postmortem" (2026-08-12).
Refs: #968, #969, SKaiNET-transformers#307.