feat(memory): MemoryPlan from the GGUF header — weights, KV, forward slab, budget fit, suggestions (SKEEP-003 P1, S0.7) - #1056
Merged
Conversation
Contributor
Author
|
Local gate Targeted: lang-core |
…boQuant), forward slab, headroom, budget fit, suggestions (SKEEP-003 P1) Milestone M0 (#1001), PRD M0-F1..F3 / M0-A4: "will this model fit on this device at this context length?" answered from shapes and encodings only — no tensor bytes are read. - sk.ainet.lang.memory.plan: PlanTensor (name, TensorId?, Format, count, bytes; allocation = AllocationSpec mapped/MODEL/read-only), ModelGeometry, KvCacheMode (BF16, TURBOQUANT_4 via TensorEncoding.TurboQuantPolar), PlanInput (model, weights, geometry, ctx, prefill chunk 256, kv mode; unmappedWeights listed, never dropped), Budget (explicit or available − reserve: 700 MB Android/JVM, 300 MB native — decision #11), MemoryPlan (weights resident · kv @ ctx with the alternate mode · forward slab · heap headroom · total · fits · ≥ 2 suggestions with savings: --kv turboquant, --ctx N/2, smaller model; render() = the PRD §4.3 table), MemoryPlans.plan(input, budget) with the documented estimates. - io-gguf: StreamingGGUFReader.planInput(ctx, prefillChunk, kvMode, nameMap) — header only (tensor table + <arch>.block_count / embedding_length / attention.head_count(_kv) / key_length / value_length / feed_forward_length / vocab_size / context_length); ggufGeometry(); ggufFormat(type, nBytes). - Tests: MemoryPlanTest (Llama-3.2-1B-like geometry: KV bf16 @2048 = 64 MiB, TurboQuant ≈ ¼, forward slab scaling, totals/residency, does-not-fit suggestions, TurboQuant mode + available-memory budget, no geometry, unmapped, byte formatting); GgufMemoryPlanTest (synthetic header-only plan, format mapping, fixture-gated Qwen2.5-0.5B real header). - BCV: lang-core jvm dump regenerated (additions only). Includes the AllocationSpec commit of #1053 (cherry-picked: the plan's line items are AllocationSpecs); that commit drops out on rebase once Closes #1012 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
michalharakal
force-pushed
the
feature/1011-gguf-namemap
branch
from
August 22, 2026 21:08
a9d4821 to
fae03eb
Compare
michalharakal
force-pushed
the
feature/1012-memory-plan
branch
from
August 22, 2026 21:08
bcad0de to
f0e2e08
Compare
This was referenced Aug 22, 2026
Contributor
Author
|
Rebased onto the updated base after #1053 merged into |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S0.7 (milestone M0 #1001, PRD M0-F1 / F2 / F3, A4):
MemoryPlan— will this model fit on this device at this context length? answered from the GGUF header only (shapes + encodings + architecture metadata); no tensor bytes are read.sk.ainet.lang.memory.plan(opt-in):PlanTensor(name,TensorId?,Format, count, bytes;allocation=AllocationSpecmapped /MODEL/ read-only),ModelGeometry,KvCacheMode { BF16, TURBOQUANT_4 }(TurboQuant bytes viaTensorEncoding.TurboQuantPolar(4, 128)),PlanInput(weights, geometry,ctx, prefill chunk 256, kv mode;unmappedWeightslisted, never dropped),Budget(explicit, oravailable − reserve: 700 MB Android/JVM, 300 MB native — decision BPETokenizer #11),MemoryPlan(weights resident · KV @ ctx with the alternate mode shown · forward slab · heap headroom · total ·fits· ≥ 2 suggestions with savings:--kv turboquant,--ctx N/2, smaller model;render()prints the PRD §4.3 table),MemoryPlans.plan(input, budget)with documented estimates (KV =layers × 2 × ctx × kvHeads × headDim × bytes; forward slab =T × (4·emb + 3·ffn + heads·ctx) × 4 B + vocab × 4 B). M1's plan-vs-actual check ([S1.9] M1: plan-vs-actual — allocation-event totals vsMemoryPlan, printed; CI assertion (> 10 % fails) #1030) calibrates the estimates against real allocations.StreamingGGUFReader.planInput(ctx, prefillChunk, kvMode, nameMap)— header only: tensor table +<arch>.block_count / embedding_length / attention.head_count(_kv) / attention.key_length / value_length / feed_forward_length / vocab_size / context_length;ggufGeometry(),ggufFormat(type, nBytes).MemoryPlanTest— Llama-3.2-1B-like geometry (KV bf16 @ 2048 = 64 MiB, TurboQuant ≈ ¼, forward slab scaling, totals/residency, does-not-fit → suggestions with savings, TurboQuant mode + available-memory budget, no-geometry, unmapped, byte formatting);GgufMemoryPlanTest— synthetic header-only plan, format mapping, fixture-gated Qwen2.5-0.5B real header (geometry 24 layers / 896 / 14 heads / 2 kv heads, weights within 2 % of the packed tensor bytes).Stacked on #1055 (S0.6);
AllocationSpec(#1053, merged intofeature/1008) is in the base — the plan's line items areAllocationSpecs. Theskainet-planCLI (#1057) only renders this.Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25); results in the first comment. Targeted: lang-coresk.ainet.lang.memory.*22/22, io-gguf plan + name-map tests 6/6.Closes #1012
🤖 Generated with Claude Code