Skip to content

feat(memory): MemoryPlan from the GGUF header — weights, KV, forward slab, budget fit, suggestions (SKEEP-003 P1, S0.7) - #1056

Merged
michalharakal merged 1 commit into
developfrom
feature/1012-memory-plan
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1012-memory-plan

Conversation

@michalharakal

@michalharakal michalharakal commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S0.7 (milestone M0 #1001, PRD M0-F1 / F2 / F3, A4): MemoryPlan — will this model fit on this device at this context length? answered from the GGUF header only (shapes + encodings + architecture metadata); no tensor bytes are read.

  • sk.ainet.lang.memory.plan (opt-in): PlanTensor (name, TensorId?, Format, count, bytes; allocation = AllocationSpec mapped / MODEL / read-only), ModelGeometry, KvCacheMode { BF16, TURBOQUANT_4 } (TurboQuant bytes via TensorEncoding.TurboQuantPolar(4, 128)), PlanInput (weights, geometry, ctx, prefill chunk 256, kv mode; unmappedWeights listed, never dropped), Budget (explicit, or available − reserve: 700 MB Android/JVM, 300 MB native — decision BPETokenizer #11), MemoryPlan (weights resident · KV @ ctx with the alternate mode shown · forward slab · heap headroom · total · fits · ≥ 2 suggestions with savings: --kv turboquant, --ctx N/2, smaller model; render() prints the PRD §4.3 table), MemoryPlans.plan(input, budget) with documented estimates (KV = layers × 2 × ctx × kvHeads × headDim × bytes; forward slab = T × (4·emb + 3·ffn + heads·ctx) × 4 B + vocab × 4 B). M1's plan-vs-actual check ([S1.9] M1: plan-vs-actual — allocation-event totals vs MemoryPlan, printed; CI assertion (> 10 % fails) #1030) calibrates the estimates against real allocations.
  • io-gguf: StreamingGGUFReader.planInput(ctx, prefillChunk, kvMode, nameMap) — header only: tensor table + <arch>.block_count / embedding_length / attention.head_count(_kv) / attention.key_length / value_length / feed_forward_length / vocab_size / context_length; ggufGeometry(), ggufFormat(type, nBytes).
  • Tests: MemoryPlanTest — Llama-3.2-1B-like geometry (KV bf16 @ 2048 = 64 MiB, TurboQuant ≈ ¼, forward slab scaling, totals/residency, does-not-fit → suggestions with savings, TurboQuant mode + available-memory budget, no-geometry, unmapped, byte formatting); GgufMemoryPlanTest — synthetic header-only plan, format mapping, fixture-gated Qwen2.5-0.5B real header (geometry 24 layers / 896 / 14 heads / 2 kv heads, weights within 2 % of the packed tensor bytes).
  • BCV: lang-core jvm dump regenerated (additions only).

Stacked on #1055 (S0.6); AllocationSpec (#1053, merged into feature/1008) is in the base — the plan's line items are AllocationSpecs. The skainet-plan CLI (#1057) only renders this.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25); results in the first comment. Targeted: lang-core sk.ainet.lang.memory.* 22/22, io-gguf plan + name-map tests 6/6.

Closes #1012

🤖 Generated with Claude Code

@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on bcad0de (stacked on #1055 + #1053 cherry-pick):

=== pr-gate: JVM tests ===                                      BUILD SUCCESSFUL in 4m 11s
=== pr-gate: apiCheck ===                                       BUILD SUCCESSFUL in 18s
=== pr-gate: verifyNpmPins jsTest wasmJsTest wasmWasiTest ===   BUILD SUCCESSFUL in 2m 37s
=== pr-gate: linuxX64Test ===                                   BUILD SUCCESSFUL in 1m 29s
=== pr-gate: assemble ===                                       BUILD SUCCESSFUL
=== pr-gate: :skainet-test:skainet-test-java:test ===           BUILD SUCCESSFUL
pr-gate: all legs passed.

Targeted: lang-core sk.ainet.lang.memory.* 22/22, io-gguf GgufMemoryPlanTest + GgufNameMapFixtureTest 6/6 (fixture cases skip when the files are absent).

…boQuant), forward slab, headroom, budget fit, suggestions (SKEEP-003 P1)

Milestone M0 (#1001), PRD M0-F1..F3 / M0-A4: "will this model fit on this
device at this context length?" answered from shapes and encodings only —
no tensor bytes are read.

- sk.ainet.lang.memory.plan: PlanTensor (name, TensorId?, Format, count,
  bytes; allocation = AllocationSpec mapped/MODEL/read-only), ModelGeometry,
  KvCacheMode (BF16, TURBOQUANT_4 via TensorEncoding.TurboQuantPolar),
  PlanInput (model, weights, geometry, ctx, prefill chunk 256, kv mode;
  unmappedWeights listed, never dropped), Budget (explicit or available −
  reserve: 700 MB Android/JVM, 300 MB native — decision #11), MemoryPlan
  (weights resident · kv @ ctx with the alternate mode · forward slab ·
  heap headroom · total · fits · ≥ 2 suggestions with savings: --kv
  turboquant, --ctx N/2, smaller model; render() = the PRD §4.3 table),
  MemoryPlans.plan(input, budget) with the documented estimates.
- io-gguf: StreamingGGUFReader.planInput(ctx, prefillChunk, kvMode,
  nameMap) — header only (tensor table + <arch>.block_count /
  embedding_length / attention.head_count(_kv) / key_length / value_length /
  feed_forward_length / vocab_size / context_length); ggufGeometry();
  ggufFormat(type, nBytes).
- Tests: MemoryPlanTest (Llama-3.2-1B-like geometry: KV bf16 @2048 = 64 MiB,
  TurboQuant ≈ ¼, forward slab scaling, totals/residency, does-not-fit
  suggestions, TurboQuant mode + available-memory budget, no geometry,
  unmapped, byte formatting); GgufMemoryPlanTest (synthetic header-only
  plan, format mapping, fixture-gated Qwen2.5-0.5B real header).
- BCV: lang-core jvm dump regenerated (additions only).

Includes the AllocationSpec commit of #1053 (cherry-picked: the plan's
line items are AllocationSpecs); that commit drops out on rebase once

Closes #1012

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Rebased onto the updated base after #1053 merged into feature/1008-format-and-encoding (lang-core API dump regenerated; no source change). The chain is now: #1051 (Format + AllocationSpec) ← #1054 (TensorId) ← #1055 (NameMap) ← #1056 (MemoryPlan) ← #1057 (skainet-plan CLI); merging each into its base keeps the next one a single-commit diff.

Base automatically changed from feature/1011-gguf-namemap to develop August 23, 2026 07:17
@michalharakal
michalharakal merged commit ca09666 into develop Aug 23, 2026
3 checks passed
@michalharakal
michalharakal deleted the feature/1012-memory-plan branch August 23, 2026 07:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S0.7] P1/M0: MemoryPlan from the GGUF header — weights, KV (bf16/TurboQuant), Forward slab, headroom, budget, fit check, suggestions

1 participant