feat(plan): price the resolved weight form, so a dequantization is a table line - #1122
Merged
Merged
Conversation
…table line Closes #1116. Slice 3 of #1109. MemoryPlans.plan summed PlanTensor.bytes — what the *file* holds. If a weight is going to be dequantized at load, that is the wrong number: a budget checked against a Q4_K tensor's packed size says "fits", and then the load quadruples it. The plan now sums residentBytes, what the resolved WeightForm will actually occupy, and keeps the stored total beside it so the difference can be reported. Where that difference shows up: - the weights line reads "re-encoded at load (+N)" instead of "Mapped, packed", and reads exactly as before when nothing was re-encoded; - the fit check fails before the load rather than the allocator failing during it; - the suggestion list names it first, because unlike --ctx or the KV mode this cost was never asked for — it is what the resolver chose when no kernel on the target could feed the stored encoding. PlanInput.resolveWeightForms(profile, capabilities) is the entry point, and skainet-plan gains --kernels all|dense. Which encodings a target can feed is not a property of the machine running the planner, so it is asked for rather than detected; and the planner deliberately does not carry its own copy of the backend's encoding table, which is the drift #1114 argued against. The default, `all`, leaves every existing plan byte-identical. The issue also asked for updated golden JSON plans for the reference GGUFs. There are none in the repository — that acceptance line referenced something never built — so there was nothing to update. Gate: scripts/pr-gate.sh — all legs passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
aharakal
approved these changes
Aug 25, 2026
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1116. Slice 3 of 5 for #1109.
The failure this closes
MemoryPlans.plansummedPlanTensor.bytes— what the file holds. When a weight is going to be dequantized at load, that is the wrong number, and wrong in the dangerous direction:The plan now sums
residentBytes— what the resolvedWeightFormwill actually occupy — and keeps the stored total beside it so the difference is reportable.Where the difference surfaces
re-encoded at load (+3.0 GB)instead ofMapped, packed. A plan with nothing re-encoded renders exactly as it did before.--ctxor the KV mode this cost was never asked for — it is what the resolver chose when no kernel on the target could feed the stored encoding. The suggestion is "a build with kernels for the stored encoding", saving exactly what the conversion costs.The CLI switch, and why it is a switch
PlanInput.resolveWeightForms(profile, capabilities)is the entry point.skainet-plangains--kernels all|dense.Two reasons it is asked for rather than detected.
skainet-plandoes not depend onskainet-backend-api, soRegistryKernelCapabilitiesis out of reach — but more to the point, the planner plans for a target device, which is not the machine running the planner. Detecting the host's kernels would answer the wrong question.The alternative — giving the planner its own table of which encodings SKaiNET's CPU backend ships kernels for — is exactly the duplicated capability table #1114 argued against, and it would drift. So
--kernels denseanswers "what does this model cost on a device with no packed kernels?", and the defaultallleaves every existing plan byte-identical.Tests
Seven, including the failure above stated directly: the same weight, the same budget,
fits == trueas stored andfits == falseonce aDequantizeTo(FP32)form is resolved. Plus Q4_K → FP32 being ~8×, the conversion total, the suggestion saving what it costs, and the rendered table both showing the conversion and not mentioning it when there is none.One acceptance item had nothing behind it
The issue asked for updated "golden JSON plans for the reference GGUFs". There are none in the repository — that line referenced something never built. Nothing to update; saying so rather than quietly dropping it.
Gate
scripts/pr-gate.sh— all legs passed. (First run failed the native leg: two of my test names contained commas, which Kotlin/Native rejects in backticked identifiers and the JVM accepts. Renamed.)🤖 Generated with Claude Code