Skip to content

feat(memory): WeightForm and a resolver that asks the target, not the caller - #1119

Merged
michalharakal merged 1 commit into
developfrom
feature/1114-weight-form-resolver
Aug 25, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1114-weight-form-resolver

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1114. Slice 1 of 5 for #1109 — the type and the decision procedure, with nothing wired to it yet.

The problem

How a weight ends up in memory is decided by three independent flags on the loader constructor:

flag decides
QuantPolicy keep quantized / dequantize to FP32 / mixed
StagingPolicy heap bytes vs mapped pages
WeightOrientation [in, out] as stored vs logical [out, in]

Nobody asks the target. Nothing consults the backend about which encodings its kernels can feed, so a format with no kernel becomes a per-call dequantization nobody declared, and a format with a good kernel can still be handed bytes in an order it does not read.

And WeightOrientation stops at the shape — it reverses dimensions and explicitly leaves the bytes alone. The order that decides whether a packed kernel reads the right blocks (#973) had no name in that vocabulary. WeightByteOrder.KERNEL_FEED gives it one.

The change

WeightForm(
    encoding:  KeepAsStored | DequantizeTo(dtype) | RequantizeTo(encoding),
    order:     AS_STORED | KERNEL_FEED,
    residency: HEAP | MAPPED,
)

WeightFormResolver.resolve(stored, profile, capabilities): WeightForm

Rules, in order:

  1. Residency from the profile alone — weightsMapped is a statement about the device; a 2 GB board cannot heap its weights whatever they encode.
  2. A dense weight is already in its only form.
  3. A kernel can feed the stored encoding → keep it, and ask only whether the kernel wants feed order. Every packed matmul kernel in the tree does, so the weight arrives in the order [973.3] A weight-transposing matmul primitive; ops.transpose on packed data becomes a loud error #1096's relayout would have produced and that relayout never fires.
  4. Nothing can feed it → the backend would otherwise dequantize on every forward pass, so doing it once at load is strictly better. But it is a real cost — FP32 is ~8× a Q4_K tensor — so PlannerProfile.strict makes it a failure with the multiplier in the message. That gives strict something to do at load time; nothing consulted it there before.

Two decisions worth flagging

Placement follows the dependency graph, not the issue text. I wrote "io-core" in #1114; that is wrong. io-core and backend-api are siblings over lang-core, and the resolver needs PlannerProfile (lang-core) and kernel capabilities (backend-api) while its output is consumed by the loader (io). lang-core is the only module all three can see. So WeightForm, KernelCapabilities and the resolver live in sk.ainet.lang.memory.plan beside PlannerProfile; the registry-backed capability implementation is in backend-api where the providers are.

There are two registries, and a test caught me consulting one. KernelRegistry holds the provider SPI — the per-encoding accessors and supports, which the eager quantized paths consult. KernelDispatch holds KernelKey-addressed view kernels, and is where KernelPacks.installReference puts the reference FP32 GEMM every target is supposed to have. My first implementation asked only KernelRegistry and reported a target carrying nothing but the reference pack as unable to multiply dense floats — false, and precisely the kind of drift between capability answers and actual dispatch that this object exists to prevent. It now asks both.

I reused KernelProvider.supports("matmul", [input, weight]) rather than building a second capability table: it already exists, and it is already the contract providers override when they ship kernels beyond the built-in accessors (the ternary packs among them). A parallel table is how the two answers diverge.

Tests

11, table-driven: every encoding × {kernel present, kernel absent} × {DESKTOP, MOBILE_2GB}, plus availability (a provider whose isAvailable() is false must not make the resolver keep an encoding nothing can compute), a backend whose packed kernels read canonical order not being handed a pointless permutation, and the point of the exercise stated as an assertion — the same file resolving to different forms on two targets, with the model author writing nothing either way.

Scope

Inert by design: nothing calls the resolver. The loader is #1115, plan pricing #1116, trace events #1117, end-to-end acceptance #1118.

API: additive, 91 lines in lang-core's dump. skainet-backend-api has no BCV configured, so nothing to dump there.

Gate

scripts/pr-gate.sh — all legs passed.

🤖 Generated with Claude Code

… caller

Closes #1114. Slice 1 of 5 for #1109.

How a weight ends up in memory was decided by three independent flags whoever
constructed the loader happened to set: QuantPolicy for dequantization,
StagingPolicy for heap vs mapped, WeightOrientation for the shape. Nothing
asked the device. A format with no kernel became a per-call dequantization
nobody declared; a format with a good kernel could still be handed bytes in
an order it does not read. And WeightOrientation only reverses dimensions —
the packed byte order that decides whether a kernel reads the right blocks
(#973) had no name in that vocabulary at all.

WeightForm names all three together, and WeightFormResolver decides them from
what the file holds, what PlannerProfile says about the device, and what the
backend's kernels can feed. Residency follows the device alone. A dense
weight asks for nothing. An encoding with a kernel keeps its encoding and
gets KERNEL_FEED order, so the per-weight relayout (#1096) has nothing to do
on first use. An encoding with no kernel is dequantized once at load instead
of once per forward pass — unless the profile is strict, which finally gives
PlannerProfile.strict something to do at load time; nothing consulted it
there before.

Placement follows the dependency graph rather than the issue text: io-core
and backend-api are siblings over lang-core, and the resolver needs
PlannerProfile and kernel capabilities while its output is consumed by the
loader. lang-core is the only module all three can see. KernelCapabilities is
declared there; the registry-backed implementation is in backend-api, where
the providers are.

That implementation consults both registries, because there are two and a
test caught it consulting one. KernelRegistry holds the provider SPI;
KernelDispatch holds KernelKey-addressed view kernels, and is where
KernelPacks.installReference puts the reference FP32 GEMM. Asking only the
first reports a target carrying the reference pack as unable to multiply
dense floats — false, and exactly the drift the object exists to avoid.

Inert by design: nothing calls the resolver yet. The loader is #1115.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1119 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 97f7362 into develop Aug 25, 2026
20 checks passed
@michalharakal
michalharakal deleted the feature/1114-weight-form-resolver branch August 25, 2026 12:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WeightForm and a resolver that asks the target, not the caller

1 participant