Skip to content

feat(plan): price the resolved weight form, so a dequantization is a table line - #1122

Merged
michalharakal merged 1 commit into
developfrom
feature/1116-plan-prices-weight-form
Aug 25, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1116-plan-prices-weight-form

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1116. Slice 3 of 5 for #1109.

The failure this closes

MemoryPlans.plan summed PlanTensor.bytes — what the file holds. When a weight is going to be dequantized at load, that is the wrong number, and wrong in the dangerous direction:

budget checked against Q4_K packed size  →  "✔ fits"
load dequantizes to FP32                 →  4× the bytes  →  OOM

The plan now sums residentBytes — what the resolved WeightForm will actually occupy — and keeps the stored total beside it so the difference is reportable.

Where the difference surfaces

  • The weights line: re-encoded at load (+3.0 GB) instead of Mapped, packed. A plan with nothing re-encoded renders exactly as it did before.
  • The fit check: fails in the planner, before the load, instead of in the allocator during it.
  • The suggestions: named first, because unlike --ctx or the KV mode this cost was never asked for — it is what the resolver chose when no kernel on the target could feed the stored encoding. The suggestion is "a build with kernels for the stored encoding", saving exactly what the conversion costs.

The CLI switch, and why it is a switch

PlanInput.resolveWeightForms(profile, capabilities) is the entry point. skainet-plan gains --kernels all|dense.

Two reasons it is asked for rather than detected. skainet-plan does not depend on skainet-backend-api, so RegistryKernelCapabilities is out of reach — but more to the point, the planner plans for a target device, which is not the machine running the planner. Detecting the host's kernels would answer the wrong question.

The alternative — giving the planner its own table of which encodings SKaiNET's CPU backend ships kernels for — is exactly the duplicated capability table #1114 argued against, and it would drift. So --kernels dense answers "what does this model cost on a device with no packed kernels?", and the default all leaves every existing plan byte-identical.

Tests

Seven, including the failure above stated directly: the same weight, the same budget, fits == true as stored and fits == false once a DequantizeTo(FP32) form is resolved. Plus Q4_K → FP32 being ~8×, the conversion total, the suggestion saving what it costs, and the rendered table both showing the conversion and not mentioning it when there is none.

One acceptance item had nothing behind it

The issue asked for updated "golden JSON plans for the reference GGUFs". There are none in the repository — that line referenced something never built. Nothing to update; saying so rather than quietly dropping it.

Gate

scripts/pr-gate.sh — all legs passed. (First run failed the native leg: two of my test names contained commas, which Kotlin/Native rejects in backticked identifiers and the JVM accepts. Renamed.)

🤖 Generated with Claude Code

…table line

Closes #1116. Slice 3 of #1109.

MemoryPlans.plan summed PlanTensor.bytes — what the *file* holds. If a weight
is going to be dequantized at load, that is the wrong number: a budget
checked against a Q4_K tensor's packed size says "fits", and then the load
quadruples it. The plan now sums residentBytes, what the resolved WeightForm
will actually occupy, and keeps the stored total beside it so the difference
can be reported.

Where that difference shows up:

- the weights line reads "re-encoded at load (+N)" instead of "Mapped,
  packed", and reads exactly as before when nothing was re-encoded;
- the fit check fails before the load rather than the allocator failing
  during it;
- the suggestion list names it first, because unlike --ctx or the KV mode
  this cost was never asked for — it is what the resolver chose when no
  kernel on the target could feed the stored encoding.

PlanInput.resolveWeightForms(profile, capabilities) is the entry point, and
skainet-plan gains --kernels all|dense. Which encodings a target can feed is
not a property of the machine running the planner, so it is asked for rather
than detected; and the planner deliberately does not carry its own copy of
the backend's encoding table, which is the drift #1114 argued against. The
default, `all`, leaves every existing plan byte-identical.

The issue also asked for updated golden JSON plans for the reference GGUFs.
There are none in the repository — that acceptance line referenced something
never built — so there was nothing to update.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1122 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal requested a review from aharakal August 25, 2026 14:26
@michalharakal
michalharakal merged commit f115fb0 into develop Aug 25, 2026
18 checks passed
@michalharakal
michalharakal deleted the feature/1116-plan-prices-weight-form branch August 25, 2026 14:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A declared form changes the plan, so its cost is visible before the load

2 participants