feat(memory): sliding-window KV — a wrapped window is a pair of views, and attention iterates it (SKEEP-003 P4, S2.3) - #1090
Merged
Conversation
…, and attention iterates it Closes #1036 (SKEEP-003 P4, S2.3, proposal §4.6 ring decision, decision #4; M2-F5, M2-A4). A ring that has wrapped holds its newest positions in two runs, and the choice has always been to copy them together before attention or to grow the buffer forever. This is the third option: hand the kernel the pair. - `WindowedKV(head, tail?)` — the window as the one or two runs it physically is. Both halves are ordinary TensorViews over the same Storage, so nothing is copied; `gather(scope, sink)` makes the copy *visible* as one traced adapter for kernels that cannot iterate a pair. - `KvCacheStore.keyWindow/valueWindow` are defaulted (the default copies once through readKeys, so a TurboQuant store needs no special case) and overridden zero-copy by `DefaultKvCacheStore`: strided views with the head stride left at `maxSeqLen * headDim`, and two of them when the run crosses the end of the ring. - `DefaultKvCacheStore(slidingWindow = true)` makes the buffer a ring: appending past capacity overwrites the oldest position, positions stay absolute, `windowStart` is the oldest one still held, and reads across the wrap return positions in order. Off by default — a full cache still throws exactly as before. - `WindowedAttention.decodeStep` computes softmax online (running maximum and denominator), which is what makes the pair workable: positions are consumed run after run and nothing has to exist as one array. The accumulator is the caller's output row, so it allocates nothing per token. `CompressedKvAttention` exposes the same windows. M2-A4 is asserted directly: a ring that wrapped several times and a cache that never wrapped, over the same four positions, produce bit-identical output (`assertContentEquals`) — and the test first asserts the window really is wrapped, so it cannot pass vacuously. The gather path agrees within 1e-5 and emits exactly two adapters priced in the trace; sixteen decode steps over a wrapped ring emit zero Allocation and zero AdapterInserted events. Gate: scripts/pr-gate.sh — all legs passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1036 · Phase P4 · Milestone M2 · PRD M2-F5, M2-A4 · Proposal §4.6 ring decision, §10 decision #4
The third option
A ring that has wrapped holds its newest positions in two runs —
[from, capacity)and[0, wrapEnd). The usual choices are to copy them together before every attention step, or to keep growing the buffer so it never wraps. This takes the third: hand the kernel the pair.WindowedKV(head, tail?)— the window as the one or two runs it physically is. Both halves are ordinaryTensorViews over the sameStorage, so nothing is copied and nothing is special-cased.gather(scope, sink)produces one contiguous view for kernels that cannot iterate a pair — and makes the copy visible as a traced adapter rather than an invisible per-token allocation.KvCacheStore.keyWindow/valueWindoware defaulted: the default copies once throughreadKeys, so every store — including TurboQuant — answers correctly without changes.DefaultKvCacheStoreoverrides them zero-copy: strided views with the head stride left atmaxSeqLen * headDim, and two of them when the run crosses the end of the ring.DefaultKvCacheStore(slidingWindow = true)makes the buffer an actual ring. Appending past capacity overwrites the oldest position instead of throwing, positions stay absolute,windowStartis the oldest one still held, and reads across the wrap return positions in order. Off by default — a full cache still throws exactly as it did.WindowedAttention.decodeStepcomputes the softmax online (running maximum and denominator, the flash-attention recurrence). That is what makes the pair workable — positions are consumed run after run, nothing has to exist as one array — and the running accumulator is the caller's output row, so the kernel allocates nothing per token.CompressedKvAttentionexposes the same windows.M2-A4, asserted rather than described
A ring that has wrapped several times and a cache that never wrapped, over the same four positions, produce bit-identical output (
assertContentEquals, not a tolerance). The test first assertswrapped == true, so it cannot pass by the window quietly being contiguous.Both sides run the same online-softmax kernel, which is what makes bit-identity the right bar here: the comparison is "does the ring's addressing change the answer", not "do two different summation orders agree".
Alongside it:
1e-5and emits exactly two adapters (keys, values), each priced in bytes in the trace;Allocationand zeroAdapterInsertedevents — the zero-copy claim, checked against the event stream rather than asserted in prose;readKeys;Gate
scripts/pr-gate.sh— all legs passed. 342 storage + memory tests green.Keeps develop green by
slidingWindowis an opt-in constructor parameter on the store, not a change toKvCacheConfig(no data-class primary-constructor change), and it defaults to today's behaviour — overflow still throws, positions are still absolute,readKeysstill returns what it always did. The window accessors are defaulted interface members, so no existingKvCacheStoreimplementation has to change;WindowedKVandWindowedAttentionare new and additive.🤖 Generated with Claude Code