Skip to content

feat(memory): KvCacheStore declares its Format (dtype + encoding) - #1079

Merged
michalharakal merged 1 commit into
developfrom
feature/1077-kv-format
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1077-kv-format

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1077 · Phase P2 · Milestone M1 · Proposal §3 (Format = (DType, TensorEncoding)), SKEEP-003 rule 3

Why

The M1 decode harness (#1032 / #1078) surfaced the drift: MemoryPlan assumed bf16 KV while DefaultKvCacheStore holds FP32, understating a dense ring by 2×. The cause is that a KV store could say "4 bytes per element" but not "FP32" — it exposed only a TensorEncoding, so nothing tied the width to a logical dtype. #1076 added KvCacheMode.FP32 as the stopgap; this makes the store the source of truth.

What

  • KvCacheStore.keyFormat / valueFormat: Format(dtype, encoding) — Format(FP32, Dense(4)) for a dense ring, Format(BF16, Dense(2)) for a narrow-float one, Format(FP32, TurboQuantPolar(…)) for a compressed one (logical dtype never erased — rule 3). Both default to Format(FP32, <existing encoding>), which is exactly what every store held before, so no implementation breaks.
  • keyBytesPerElement / valueBytesPerElement — the (possibly fractional) width, derived from Format.physicalBytes so packed encodings report their true bpw rather than the dtype size.
  • KvCacheConfig.keyDType / valueDType (default FP32) — a narrow-float ring is now declarable; DefaultKvCacheStore reports the configured format, TurboQuantKvCacheStore its packed one.
  • plan.kvBytesFor(format, elements) — the width the planner should read from the store instead of guessing through KvCacheMode.

Why not generics

KvCacheStore<T : DType, V> was considered and rejected: it pushes the storage representation onto every caller (attention, sdpa, the transformers stack), and a compressed store could not honestly satisfy V = Float — it stores 4-bit polar codes and decodes on read. The FloatArray boundary is a decoded-value contract (SKEEP-003 rule 4: get() decodes, never a raw byte), with encode/decode inside the store. Format is how the rest of the architecture already describes what it is and how it is stored, so the cache now speaks the same language as TensorView, AllocationSpec and MemoryPlan.

Acceptance

KvCacheFormatTest (new):

181/181 sk.ainet.lang.tensor.storage.* tests pass.

Gate

scripts/pr-gate.sh — all legs passed: jvmTest · apiCheck (dumps regenerated, additive only) · verifyNpmPins jsTest wasmJsTest wasmWasiTest · linuxX64Test · assemble · :skainet-test:skainet-test-java:test.

Keeps develop green by

Every new member is a defaulted interface property — existing KvCacheStore implementations (in-repo and downstream) compile unchanged and report the FP32 format they already had. KvCacheConfig's new dtype parameters default to FP32, so existing construction sites are unaffected. kvBytesFor is additive; no planner behaviour changes in this PR (wiring the planner to read a supplied store's format is the follow-up now that the store can answer).

🤖 Generated with Claude Code

…tead of only an encoding

Closes #1077, follow-up to the drift the M1 harness found (#1078): the
planner assumed bf16 KV while DefaultKvCacheStore holds FP32, understating
a dense ring by 2x. The cache could say "4 bytes per element" but not
"FP32", so nothing connected the width to a dtype.

- KvCacheStore.keyFormat / valueFormat: Format(dtype, encoding) —
  Format(FP32, Dense(4)) for a dense ring, Format(BF16, Dense(2)) for a
  narrow-float one, Format(FP32, TurboQuantPolar(4, 128)) for a
  compressed one. Both default to Format(FP32, <encoding>), which is what
  every store held before, so no implementation breaks; keyBytesPerElement
  / valueBytesPerElement give the (possibly fractional) width.
- KvCacheConfig gains keyDType / valueDType (default FP32), so a
  narrow-float ring is declarable; DefaultKvCacheStore reports the
  configured format and TurboQuantKvCacheStore its packed one.
- plan.kvBytesFor(format, elements): the width the planner should use —
  read from the store rather than guessed through KvCacheMode.

Not generics: KvCacheStore<T : DType, V> would push the storage
representation onto every caller (attention, sdpa, the transformers
stack) and a compressed store could not honestly satisfy V = Float. The
FloatArray boundary is a decoded-value contract (SKEEP-003 rule 4) with
encode/decode inside the store; Format is how the rest of the
architecture already describes "what it is and how it is stored".

KvCacheFormatTest: the dense store declares FP32; a narrow-float ring is
declarable and the planner's width follows the declaration; a compressed
store reports its packed format and a sub-FP32 width; a dense ring is
planned FP32-wide, not bf16. 181/181 storage tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 52f59ca into develop Aug 23, 2026
14 checks passed
@michalharakal
michalharakal deleted the feature/1077-kv-format branch August 23, 2026 20:41
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1079 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

KvCacheStore should declare its Format (dtype + encoding), not just an encoding

1 participant