Skip to content

feat(memory): KV cache preallocates its ring in Scope.Model (SKEEP-003 P2, S1.10) - #1076

Merged
michalharakal merged 1 commit into
developfrom
feature/1031-kv-model-scope-prealloc
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1031-kv-model-scope-prealloc

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.10 (milestone M1 #1002, PRD M1-F2): the KV cache preallocates its ring in Scope.Model.

The cache already preallocated the whole ring at construction — what it lacked was an owner. DefaultKvCacheStore now takes an optional ModelScope and allocates the per-layer K/V backing through it:

  • tracked and traced — one Allocation event per layer and side, in MODEL scope, carrying the cache's TensorId (kv.layers[3].k / .v) rather than an anonymous buffer, so the ring shows up in plan-vs-actual (feat(memory): plan-vs-actual — reconstruct a run's memory from the event stream and fail on drift (SKEEP-003 P2, S1.9) #1074) as model-scope bytes and never as forward-scope churn;
  • released deterministically when the model closes (Free events summing exactly to the ring size), instead of waiting for the GC — the Model.load(...).use { } story of §8 item 2;
  • preallocatedBytes exposes what the ring costs (layers × 2 × heads × maxSeqLen × headDim × 4), so a caller can compare it with the plan's KV line directly.

Without a scope the store behaves exactly as before (plain arrays, GC lifetime), so nothing existing changes. The TurboQuant store keeps its block-encoded path; giving it the same owner is a follow-up.

Evidence (KvCacheModelScopeTest, 177/177 storage tests): the no-scope path is unchanged; with a scope the ring is allocated once per layer/side in MODEL scope with the right ids and total; appending tokens allocates nothing more; closing the model frees exactly the ring; and ActualMemory sees it as model-scope peak with zero forward bytes — the weights-vs-activations separation #1065 exists to enforce.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) — all legs passed; results in the first comment.

Closes #1031

🤖 Generated with Claude Code

…3 P2)

Milestone M1 (#1002), PRD M1-F2. The cache already preallocated its whole
ring; what it lacked was an owner. DefaultKvCacheStore now takes an
optional ModelScope and allocates the per-layer K/V backing through it:

- tracked and traced — one Allocation event per layer and side, in MODEL
  scope, carrying the cache's TensorId (kv.layers[N].k / .v) instead of an
  anonymous buffer, so the ring shows up in plan-vs-actual (#1074) as
  model-scope bytes and never as forward-scope churn;
- released deterministically when the model closes (Free events summing
  to the ring size), instead of waiting for the GC;
- preallocatedBytes exposes what the ring costs
  (layers × 2 × heads × maxSeqLen × headDim × 4) so a caller can compare
  it with the plan's KV line.

Without a scope the store behaves exactly as before (plain arrays, GC
lifetime), so nothing existing changes. The TurboQuant store keeps its
block-encoded path; giving it the same owner is a follow-up.

KvCacheModelScopeTest: the no-scope path is unchanged; with a scope the
ring is allocated once per layer/side in MODEL scope with the right ids
and total, appending tokens allocates nothing more, closing the model
frees exactly the ring, and ActualMemory sees it as model-scope peak with
zero forward bytes. 177/177 storage tests.

Closes #1031

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on 4cf5096a: all legs passed — jvmTest · apiCheck · JS/Wasm · linuxX64Test · assemble · Java consumer API tests.

Targeted: sk.ainet.lang.tensor.storage.* 177/177.

@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1076 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 6d229a5 into develop Aug 23, 2026
17 checks passed
@michalharakal
michalharakal deleted the feature/1031-kv-model-scope-prealloc branch August 23, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.10] P2: KV cache preallocated in Scope.Model to ctx (KvCacheStore), TurboQuant store included

1 participant