Skip to content

feat(memory): sliding-window KV — a wrapped window is a pair of views, and attention iterates it (SKEEP-003 P4, S2.3) - #1090

Merged
michalharakal merged 1 commit into
developfrom
feature/1036-sliding-window-kv-sdpa
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1036-sliding-window-kv-sdpa

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1036 · Phase P4 · Milestone M2 · PRD M2-F5, M2-A4 · Proposal §4.6 ring decision, §10 decision #4

The third option

A ring that has wrapped holds its newest positions in two runs — [from, capacity) and [0, wrapEnd). The usual choices are to copy them together before every attention step, or to keep growing the buffer so it never wraps. This takes the third: hand the kernel the pair.

  • WindowedKV(head, tail?) — the window as the one or two runs it physically is. Both halves are ordinary TensorViews over the same Storage, so nothing is copied and nothing is special-cased. gather(scope, sink) produces one contiguous view for kernels that cannot iterate a pair — and makes the copy visible as a traced adapter rather than an invisible per-token allocation.
  • KvCacheStore.keyWindow / valueWindow are defaulted: the default copies once through readKeys, so every store — including TurboQuant — answers correctly without changes. DefaultKvCacheStore overrides them zero-copy: strided views with the head stride left at maxSeqLen * headDim, and two of them when the run crosses the end of the ring.
  • DefaultKvCacheStore(slidingWindow = true) makes the buffer an actual ring. Appending past capacity overwrites the oldest position instead of throwing, positions stay absolute, windowStart is the oldest one still held, and reads across the wrap return positions in order. Off by default — a full cache still throws exactly as it did.
  • WindowedAttention.decodeStep computes the softmax online (running maximum and denominator, the flash-attention recurrence). That is what makes the pair workable — positions are consumed run after run, nothing has to exist as one array — and the running accumulator is the caller's output row, so the kernel allocates nothing per token. CompressedKvAttention exposes the same windows.

M2-A4, asserted rather than described

A ring that has wrapped several times and a cache that never wrapped, over the same four positions, produce bit-identical output (assertContentEquals, not a tolerance). The test first asserts wrapped == true, so it cannot pass by the window quietly being contiguous.

Both sides run the same online-softmax kernel, which is what makes bit-identity the right bar here: the comparison is "does the ring's addressing change the answer", not "do two different summation orders agree".

Alongside it:

  • the gather path agrees with the pair path within 1e-5 and emits exactly two adapters (keys, values), each priced in bytes in the trace;
  • sixteen decode steps over a wrapped ring emit zero Allocation and zero AdapterInserted events — the zero-copy claim, checked against the event stream rather than asserted in prose;
  • positions come back in order across the wrap, both through the window and through readKeys;
  • a ring refuses windows that start before the oldest position it still holds, instead of silently returning overwritten data.

Gate

scripts/pr-gate.sh — all legs passed. 342 storage + memory tests green.

Keeps develop green by

slidingWindow is an opt-in constructor parameter on the store, not a change to KvCacheConfig (no data-class primary-constructor change), and it defaults to today's behaviour — overflow still throws, positions are still absolute, readKeys still returns what it always did. The window accessors are defaulted interface members, so no existing KvCacheStore implementation has to change; WindowedKV and WindowedAttention are new and additive.

🤖 Generated with Claude Code

…, and attention iterates it

Closes #1036 (SKEEP-003 P4, S2.3, proposal §4.6 ring decision, decision #4; M2-F5, M2-A4).

A ring that has wrapped holds its newest positions in two runs, and the
choice has always been to copy them together before attention or to grow
the buffer forever. This is the third option: hand the kernel the pair.

- `WindowedKV(head, tail?)` — the window as the one or two runs it
  physically is. Both halves are ordinary TensorViews over the same
  Storage, so nothing is copied; `gather(scope, sink)` makes the copy
  *visible* as one traced adapter for kernels that cannot iterate a pair.
- `KvCacheStore.keyWindow/valueWindow` are defaulted (the default copies
  once through readKeys, so a TurboQuant store needs no special case) and
  overridden zero-copy by `DefaultKvCacheStore`: strided views with the
  head stride left at `maxSeqLen * headDim`, and two of them when the run
  crosses the end of the ring.
- `DefaultKvCacheStore(slidingWindow = true)` makes the buffer a ring:
  appending past capacity overwrites the oldest position, positions stay
  absolute, `windowStart` is the oldest one still held, and reads across
  the wrap return positions in order. Off by default — a full cache still
  throws exactly as before.
- `WindowedAttention.decodeStep` computes softmax online (running maximum
  and denominator), which is what makes the pair workable: positions are
  consumed run after run and nothing has to exist as one array. The
  accumulator is the caller's output row, so it allocates nothing per
  token. `CompressedKvAttention` exposes the same windows.

M2-A4 is asserted directly: a ring that wrapped several times and a cache
that never wrapped, over the same four positions, produce bit-identical
output (`assertContentEquals`) — and the test first asserts the window
really is wrapped, so it cannot pass vacuously. The gather path agrees
within 1e-5 and emits exactly two adapters priced in the trace; sixteen
decode steps over a wrapped ring emit zero Allocation and zero
AdapterInserted events.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1090 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 9913094 into develop Aug 24, 2026
17 checks passed
@michalharakal
michalharakal deleted the feature/1036-sliding-window-kv-sdpa branch August 24, 2026 12:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S2.3] P4: sliding-window KV — window(from,to) → (head, tail) pair accepted by sdpa; gather adapter fallback; ring wrap-around parity

1 participant