Skip to content

feat(io): weight orientation at the load boundary, and a guard that refuses a transposed label (#973.5) - #1103

Merged
michalharakal merged 1 commit into
developfrom
feature/1098-weight-orientation
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1098-weight-orientation

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1098 · #973 census contradiction #6

The mismatch

GGUF writes dimensions in ne order — fastest-varying first — so a weight the engine calls [out, in] arrives from the streaming loader labelled [in, out], while its bytes are already [out, in] row-major. Only the label is wrong.

The label is what the block relayout reads. Driven by [in, out] it permutes the wrong grid, or require-fails because out is not a multiple of the block size. Both failures are named in the census; neither was detectable from the value.

Two changes, both opt-in

WeightOrientation { AS_STORED, OUT_IN } on StreamingGgufParametersLoader. OUT_IN reverses 2-D weights at the boundary and touches nothing else — not the bytes, and not 1-D tensors, which have no orientation to get wrong. It defaults to AS_STORED (today's behaviour) because reversing shapes changes what every consumer sees; new code should ask for OUT_IN, and making it the default is a follow-up once downstream has moved.

PackedWeights.requireOutIn(rows, inputDim, encoding), run by prepackForMatmul: refuses a weight that looks transposed rather than computing a wrong permutation from it, with a message that names the fix.

It is a heuristic and says so in its own kdoc. It fires when the first dimension is block-aligned and the second is not — exactly what ne order produces for a real weight — and stays quiet when both are aligned and it genuinely cannot tell. A false refusal on an ambiguous shape would be worse than no check, so the guard is deliberately one-sided.

Something the census implies and the doc now records

The two GGUF readers in this repository disagree about this today: the legacy GGUFReader reverses dimensions (npDims = dims.reversed()), the streaming reader does not. That is a live inconsistency in one module, and WeightOrientation is how a caller states which behaviour it wants instead of discovering it.

Acceptance

  • SyntheticGguf can now write multi-dimensional tensors, so the test writes a real ne = [64, 4] weight: AS_STORED reports [64, 4], OUT_IN reports [4, 64], and the decoded values are identical either way — the claim that only the label changes.
  • A 1-D tensor is unaffected by either setting.
  • The guard refuses [128, 3] Q8_0 with a message containing both "looks like [in, out]" and the WeightOrientation.OUT_IN fix, accepts [3, 128], stays quiet on [64, 128] and [128, 128], refuses an input dimension that is unaligned whichever way round it is, and ignores dense encodings entirely.

Gate

scripts/pr-gate.sh — all legs passed.

Keeps develop green by

A new enum with a default equal to today's behaviour, and a guard that only fires on the shape that was already broken. No existing call site changes.

#973 after this

#1094 ✅ block order in the type system + normative doc
#1095 ✅ packed SPI kernels bridged through the ordered key
#1097 ✅ engine-owned relayout + cross-repo fixtures
#1098 ✅ orientation at the load boundary (this PR)
#1096 matmulWT primitive; packed ops.transpose becomes a loud error

#1096 is the only structural change left — and the only one that needs downstream coordination, since it deprecates a public op and removes the per-forward O(bytes) copy Linear.onForward currently pays.

🤖 Generated with Claude Code

…efuses a transposed label

Closes #1098 (#973 census contradiction #6).

GGUF writes dimensions in `ne` order, so a weight the engine calls
`[out, in]` arrives from the streaming loader labelled `[in, out]` — while
its bytes are already `[out, in]` row-major. Only the label is wrong, and
the label is what the block relayout reads: driven by `[in, out]` it
permutes the wrong grid, or refuses because `out` is not a multiple of the
block size. Both failures are in the census.

- `WeightOrientation { AS_STORED, OUT_IN }` on
  `StreamingGgufParametersLoader`. `OUT_IN` reverses 2-D weights at the
  boundary and touches nothing else — not the bytes, not 1-D tensors, which
  have no orientation to get wrong. Defaults to `AS_STORED`, today's
  behaviour, because reversing shapes changes what every consumer sees.
- `PackedWeights.requireOutIn(rows, inputDim, encoding)`, run by
  `prepackForMatmul`: refuses a weight that looks transposed instead of
  computing a wrong permutation from it, and names the fix. It is a
  heuristic and says so — it fires when the *first* dimension is
  block-aligned and the second is not, which is exactly what `ne` order
  produces, and stays quiet when both are aligned and it cannot tell. A
  false refusal would be worse than none.
- The normative doc gains the orientation section, including the fact that
  this repository's two GGUF readers disagree about it today: the legacy
  `GGUFReader` reverses dimensions, the streaming one does not.
  `WeightOrientation` is how a caller states which it wants.

`SyntheticGguf` can now write multi-dimensional tensors, so the test writes
a real `ne = [64, 4]` weight and asserts that `AS_STORED` reports
`[64, 4]`, `OUT_IN` reports `[4, 64]`, and the decoded values are identical
either way.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1103 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 3129fe3 into develop Aug 24, 2026
19 checks passed
@michalharakal
michalharakal deleted the feature/1098-weight-orientation branch August 24, 2026 19:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[973.5] Normalize weight orientation at the load boundary — every 2-D matmul weight enters as [out, in]

1 participant