Skip to content

fix(quant): stop pre-relayouting packed weights, bump engine to 0.40.1 - #311

Merged
michalharakal merged 1 commit into
developfrom
bump/skainet-0.40.1
Aug 12, 2026
Merged

michalharakal merged 1 commit into
developfrom
bump/skainet-0.40.1

Conversation

@michalharakal

@michalharakal michalharakal commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Bumps the skainet engine pin from 0.40.0 to 0.40.1 (SKaiNET#968 hotfix, PR SKaiNET-developers/SKaiNET#969) and fixes a regression that pin bump exposes in the classic packed-quant path.

Root cause: Engine 0.40.1 changed ops.transpose on packed block-quant tensors from a shape-only relabel into a physical canonical→kernel-native block-grid permutation, fixing SK#968. BlockQuantPacking.pack() here was eagerly relayouting checkpoint bytes into kernel-native order at load time and relying on the (formerly no-op) engine transpose to pass them through unchanged at forward time. On 0.40.1 that transpose now double-permutes the block grid, silently producing wrong matmul output on the classic (non-pre-transposed) path. The pre-transposed production path (packPreTransposed, what gemma/llama/kgemma actually ship) never calls ops.transpose() and is unaffected — this only hits the classic/fallback path and, separately, an Apertus converter bug described below.

Full writeup: see packed-quant-layout-postmortem.md (SKaiNET engine repo, this session) for the cross-repo layout-convention analysis this fix is scoped from — the atomic "immediate unblock" item (§6.1) in that doc's remediation plan.

This PR is the tactical fix only. The underlying root cause — packed-quant byte order is an unwritten, contradictory contract across the engine and its converters — is tracked separately at SKaiNET-developers/SKaiNET#973, since it lives mostly in the engine repo and needs a structural fix (not another byte-order guess in transpose()), not another patch here.

Changes

  • BlockQuantPacking.pack(): stores checkpoint bytes verbatim (canonical, [out,in]) instead of relayouting, matching what the 0.40.1 ops.transpose now expects on its input.
  • ApertusMemSegConverter: its Q4_K/Q6_K path had its own inlined relayout under an unmarked [out,in] shape — it never got migrated to packPreTransposed like gemma/llama were (the Hoist quant packing + RowDequantSource out of sk.ainet.models.gemma into shared layers #184 hoist missed it). That inlined path is corrupted by 0.40.1 independent of this bug's mechanism (it relied on the old no-op transpose too). Switched to packPreTransposed, aligning it with the other converters.
  • New tests: a parity matrix test covering all 7 packed formats through both the classic and pre-transposed paths against an FP32 reference (LinearProjectionPackedParityMatrixTest), a synthetic Apertus converter test that doesn't need a real checkpoint (ApertusMemSegConverterSyntheticTest), and a llama quant-layout test (LlamaQuantLayoutTest) — closing the coverage gap that let this regression through (previously only Q5_1 was covered end-to-end).

Test plan

  • :transformer-core:jvmTest --tests '*LinearProjectionPreTransposedTest*' — the original repro from the handoff; A/B'd against 0.40.0 (fails on 0.40.1 without this fix, passes with it)
  • :transformer-core:jvmTest (full module, incl. new parity matrix + BlockQuantPackingTest) — passes
  • :llm-inference:apertus:jvmTest (incl. new synthetic test) — passes
  • :llm-inference:gemma:jvmTest — passes
  • :llm-inference:llama:jvmTest (incl. new LlamaQuantLayoutTest) — passes
  • Real-checkpoint gemma integration tests (-PincludeIntegration, functiongemma-physical-ai-v10-Q5_K_M.gguf, 260MB, FunctionGemma-270M Q5_K_M): GemmaDslQ4KTest, GemmaQ5xPackedParityTest (3 tests incl. both previously-failing q5_0/q5_1 packed-linearProject cases), GemmaQ5KPackedParityTest (production pre-transposed path, 122.5s) — all pass, 0 skips, 0 failures. This closes the last open gap; every failure from the original regression handoff is now confirmed fixed.

All test plan items are now verified. Ready for human review.

🤖 Generated with Claude Code

Engine 0.40.1 (SKaiNET#968 / PR #969) changed ops.transpose on packed
block-quant tensors from a shape-only relabel into a physical
canonical->kernel-native block-grid permutation. BlockQuantPacking.pack()
was eagerly relayouting checkpoint bytes into kernel-native order and
relying on the (formerly no-op) transpose to pass them through unchanged,
so on 0.40.1 the classic packed path double-permutes the block grid and
silently produces wrong matmul output. The pre-transposed production path
(packPreTransposed) never calls transpose() and is unaffected.

- BlockQuantPacking.pack() now stores checkpoint bytes verbatim
  (canonical, [out,in]) instead of relayouting, matching what the 0.40.1
  transpose expects on its classic-path input.
- ApertusMemSegConverter's Q4_K/Q6_K path had its own inlined relayout
  under an unmarked [out,in] shape (never migrated to packPreTransposed
  like gemma/llama). Switched it to packPreTransposed, which also fixes
  it under 0.40.1.
- Added a parity matrix test covering all 7 packed formats through both
  the classic and pre-transposed paths against an FP32 reference, a
  synthetic Apertus converter test, and a llama quant-layout test, to
  close the coverage gap that let this regression through undetected.

Verified: transformer-core, apertus, gemma, and llama module test suites
pass on 0.40.1. Real-checkpoint gemma integration tests (gated behind
GEMMA_GGUF) not run in this environment — checkpoint unavailable locally.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@michalharakal
michalharakal marked this pull request as ready for review August 12, 2026 15:01
@michalharakal
michalharakal merged commit e24971c into develop Aug 12, 2026
2 checks passed
@michalharakal
michalharakal deleted the bump/skainet-0.40.1 branch August 12, 2026 15:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant