Skip to content

Declared in-memory weight form, resolved from the target and the file (repack and dequant as stated intent) #1109

Description

@michalharakal

Follows #1108, #973, and the M0 planner work (#1001).

The problem

How a weight ends up in memory is currently decided by three independent flags the caller sets on the loader:

existing in decides
StagingPolicy skainet-io-core heap bytes vs mapped pages
QuantPolicy skainet-io-core keep quantized / dequantize to FP32 / mixed
WeightOrientation skainet-io-core [in, out] as stored vs logical [out, in]

Three problems with that:

  1. Nobody consults the target. Nothing asks the backend what its kernels can actually feed on. A Q4_K weight loaded on a target with no Q4_K kernel is a per-call dequantization nobody declared; the same weight on a target with a NEON kernel wants input-block-major bytes and gets canonical ones.
  2. WeightOrientation stops at the shape. It reverses dimensions; it does not touch packed byte order. The order that matters to a kernel (Packed-quant byte-order (block layout) is an unwritten, contradictory contract across the engine and downstream converters #973) has no name in this vocabulary.
  3. It is a caller decision that should be a resolved one. The information needed — what the file holds, what the device has, what the plan can afford — is available at load time, in one place. Making the user pick from three enums is asking them to do the resolution by hand, for a target they may not be building for.

Proposal

One declared desired in-memory form per model (overridable per tensor), resolved at load:

WeightForm(
    encoding: KeepAsStored | DequantizeTo(DType) | RequantizeTo(TensorEncoding),
    order:    AS_STORED | KERNEL_FEED,
    residency: HEAP | MAPPED,          // today's StagingPolicy
)

with a resolver whose inputs are what the file holds + PlannerProfile + the backend's kernel capabilities (already expressed in KernelKey's capability set). The default is "whatever the best available kernel for this encoding wants" — so the ordinary user declares nothing, and the answer differs correctly between a desktop and a 2 GB board.

Repacking and dequantizing become stated intent, executed once at load, rather than accidents discovered per forward pass. Concretely:

Which is the point of tying it to the planner: MemoryPlan already models bytes per tensor, so a declared form changes the plan, and skainet-plan <gguf> shows the cost before the load. A form that does not fit the budget is a planner failure with a suggestion, not an OOM. Each conversion emits a TraceSink adapter event, so "why is this model 3 GB" has an answer in the trace.

Relationship to #1108

Complementary, not alternative. #1108 makes the DSL work whatever form the weight is in. This makes the form a decision someone made on purpose, with the cost visible. With both, the common path is: the loader put the weight in kernel feed order, so #1108's marker resolves to a matmul with no relayout at all.

Acceptance

  • WeightForm + resolver in skainet-io-core, defaulted from PlannerProfile × backend kernel capabilities
  • StagingPolicy / QuantPolicy / WeightOrientation expressed through it, deprecated with ReplaceWith, behaviour unchanged for callers who set them explicitly
  • A repack or dequant chosen by the resolver appears in the MemoryPlan and in skainet-plan output, before the load
  • Each conversion emits a trace event naming the tensor, the from-form and the to-form
  • Golden parity: for every encoding, a model loaded under a resolved form produces bit-identical output to the same model loaded AS_STORED
  • Same model, same code, two profiles (DESKTOP / MOBILE_2GB) resolve to different forms and both run

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions