Skip to content

feat(io): the GGUF loader takes a WeightForm; the three flags become views of it - #1121

Merged
michalharakal merged 1 commit into
developfrom
feature/1115-loader-weight-form
Aug 25, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1115-loader-weight-form

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1115. Slice 2 of 5 for #1109.

quantPolicy, staging and weightOrientation are three constructor parameters asking the caller for three axes of one decision — about a device they may not be building for. The loader now takes a WeightForm and derives the three from it.

StreamingGgufParametersLoader(sourceProvider = …, weightForm = form)

weightForm defaults to null, meaning "use the three parameters", so every existing caller is byte-identical. Passing both a form and a flag is refused, not silently resolved: one of them would have to lose, and a caller who set staging = MAPPED believes it is in effect.

Two corrections to what the issues said

#1114's WeightForm was missing an axis. WeightOrientation reverses the shape label and explicitly leaves bytes alone. WeightByteOrder is about block placement. Those are orthogonal — reversing a packed weight's dimensions moves no data, while changing its block order moves every block and leaves the shape alone — and slice 1 collapsed them into one. WeightForm gains shape: WeightShapeOrientation, so the three old flags map cleanly onto three of its four axes and order is the genuinely new one.

KERNEL_FEED is split out to #1120 and refused here. Producing feed-order bytes at load is the easy part — TensorView.prepack(INPUT_BLOCK_MAJOR) already does exactly that permutation, in lang-core, and already emits the trace event #1117 wants. The problem is what the bytes then claim to be: Q4_KBlockTensorData and friends address packedData as canonical row-major in getBlockScale, getCode and dequantizeBlock, so feed-order bytes would decode the wrong elements and not fail. That is #973 and #968 arrived at from the loader side instead of the transpose side.

So the loader refuses it, and the message says why and where it is being fixed rather than "unsupported":

WeightByteOrder.KERNEL_FEED is not supported by this loader yet (#1120). The bytes are easy — TensorView.prepack(INPUT_BLOCK_MAJOR) does the permutation — but packed TensorData addresses packedData as canonical row-major …

EncodingRequest.RequantizeTo and DequantizeTo(non-FP32) are refused for their own reasons, up front rather than mid-load.

Deprecations

StagingPolicy and WeightOrientation get a real ReplaceWith — same two values each, drop-in. QuantPolicy gets a message only: EncodingRequest is not a drop-in for it (NATIVE_OPTIMIZED → KeepAsStored, DEQUANTIZE_TO_FP32 → DequantizeTo(FP32), and RAW_BYTES has no counterpart because no loader ever supported it). A ReplaceWith that does not compile is worse than none, so the issue's "deprecate with ReplaceWith" is only honoured where it can be honest.

Warnings are not errors in this build, so the deprecations do not break anything downstream.

Tests

The claim worth testing is that this is a change of spelling only. So WeightFormLoaderParityTest loads a mixed-encoding file — F32, Q4_K, a 2-D Q8_0 weight three blocks wide, F16 — under all eight combinations of the three flags, twice: once through the flags and once through the WeightForm they map to, compared on shapes and every element. A wrong mapping anywhere makes some cell of that product disagree.

Plus: the default loader equals the default form, both-set is refused, and the two refusals above assert on why they refused, not just that they did.

API

WeightForm's constructor signature changes — binary-incompatible for a class added yesterday in #1119. It is @ExperimentalMemoryApi and unreleased, so nothing can be depending on it yet; flagging it rather than letting the dump diff speak for itself.

Gate

scripts/pr-gate.sh — all legs passed.

🤖 Generated with Claude Code

…views of it

Closes #1115. Slice 2 of #1109.

quantPolicy, staging and weightOrientation are three constructor parameters
asking the caller for three axes of one decision, about a device they may not
be building for. The loader now takes a WeightForm instead, and derives the
three from it. Passing a form defaults to null — meaning "use the three
parameters" — so every existing caller is byte-identical, and passing both a
form and a flag is refused rather than silently resolved, because a caller
who set staging = MAPPED believes it is in effect.

Two corrections to what #1114 shipped and #1115 described.

WeightForm was missing an axis. WeightOrientation reverses the *shape label*
and explicitly leaves bytes alone; WeightByteOrder is about *block
placement*. They are orthogonal — a weight can need either, both or neither —
and slice 1 collapsed them. WeightForm gains `shape`, so the three old flags
map onto three of its four axes and `order` is the genuinely new one.

KERNEL_FEED is split out to #1120 and refused here. Producing feed-order
bytes is the easy part: TensorView.prepack(INPUT_BLOCK_MAJOR) already does
that permutation. The problem is what the bytes then claim to be —
Q*BlockTensorData addresses packedData as canonical row-major in
getBlockScale, getCode and dequantizeBlock, so feed-order bytes would decode
the wrong elements without failing. That is #973/#968 reached from the loader
side, so the refusal names the reason and the issue rather than saying
"unsupported".

StagingPolicy and WeightOrientation are deprecated with a real ReplaceWith —
same values, drop-in. QuantPolicy gets a message only: EncodingRequest is not
a drop-in for it, because RAW_BYTES has no counterpart, no loader having
supported it. A ReplaceWith that does not compile is worse than none.

WeightForm's constructor signature changes, which is binary-incompatible for
a class added yesterday in #1119. It is @ExperimentalMemoryApi and unreleased.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1121 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal requested a review from aharakal August 25, 2026 14:27
@michalharakal
michalharakal merged commit b0cf47a into develop Aug 25, 2026
18 checks passed
@michalharakal
michalharakal deleted the feature/1115-loader-weight-form branch August 25, 2026 14:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The GGUF loader takes a WeightForm; the three policy enums become views of it

2 participants