You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up from the M1 acceptance work (#1032): writing the plan-vs-actual assertion surfaced that the KV cache's dtype is not declarable, only its encoding.
What is there today
KvCacheConfig carries keyEncoding / valueEncoding: TensorEncoding, defaulting to Dense(4); DefaultKvCacheStore hard-codes FloatArray rings. So a store can say "4 bytes per element" but not "FP32", and nothing connects the two.
The consequence, found by the harness: the memory planner assumed bf16 KV while the dense store holds FP32 — a 2× understatement of the ring. #1032 works around it by adding KvCacheMode.FP32 and having callers pick the right mode; the planner still guesses rather than asks.
Generic type parameters would push the storage representation onto every caller: attention code, sdpa and the whole SKaiNET-transformers stack would need to name the cache's value type just to hold a reference, and a compressed store could not honestly satisfy V = Float. The boundary API (appendToken(key: FloatArray), readKeys(): FloatArray) is a decoded value contract — SKEEP-003 rule 4 — with encode/decode inside the store. That is exactly the erasure the storage rework removes elsewhere; re-introducing it in the KV API would be a step back.
Scope
keyFormat/valueFormat on KvCacheStore (default-derived from the encodings, so no implementation breaks);
optional keyFormat/valueFormat on KvCacheConfig;
DefaultKvCacheStore honours a narrow-float format (bf16/fp16 rings) as well as FP32;
MemoryPlan takes the KV byte width from the store's format when one is supplied, instead of KvCacheMode;
TurboQuantKvCacheStore reports its TurboQuant format.
Follow-up from the M1 acceptance work (#1032): writing the plan-vs-actual assertion surfaced that the KV cache's dtype is not declarable, only its encoding.
What is there today
KvCacheConfigcarrieskeyEncoding/valueEncoding: TensorEncoding, defaulting toDense(4);DefaultKvCacheStorehard-codesFloatArrayrings. So a store can say "4 bytes per element" but not "FP32", and nothing connects the two.The consequence, found by the harness: the memory planner assumed bf16 KV while the dense store holds FP32 — a 2× understatement of the ring. #1032 works around it by adding
KvCacheMode.FP32and having callers pick the right mode; the planner still guesses rather than asks.Proposal: a
Formatper side, not type parametersFormatwas introduced for (feat(memory): Format(dtype, encoding), TensorData.encoding, Tensor/TensorStorage.format (SKEEP-003 P1, S0.4) #1051, SKEEP-003 §0).store.keyFormat.physicalBytes(elements)) instead of guessing, so the FP32-vs-bf16 drift becomes impossible rather than caught after the fact by feat(memory): plan-vs-actual — reconstruct a run's memory from the event stream and fail on drift (SKEEP-003 P2, S1.9) #1074.KernelKeylike every other operand (feat(backend-api): KernelKey dispatch — declared formats/layouts, rank normalised once, visible adapters, reference matmul (SKEEP-003 P3, S1.7a) #1070).keyEncodingkeeps working;keyFormatdefaults toFormat(FP32, keyEncoding);KvCacheConfiggains optional format parameters.DefaultKvCacheStorecan actually store bf16 when asked, which is the M2 TurboQuant-KV story ([S2.3] P4: sliding-window KV —window(from,to)→(head, tail)pair accepted by sdpa; gather adapter fallback; ring wrap-around parity #1036, [S2.6] P5: planner 2 GB reference profile — defaults, fit check, TurboQuant KV auto ≥ 80 %, dequant warn/strict, desktop profile #1039).Why not
KvCacheStore<T : DType, V>Generic type parameters would push the storage representation onto every caller: attention code, sdpa and the whole SKaiNET-transformers stack would need to name the cache's value type just to hold a reference, and a compressed store could not honestly satisfy
V = Float. The boundary API (appendToken(key: FloatArray),readKeys(): FloatArray) is a decoded value contract — SKEEP-003 rule 4 — with encode/decode inside the store. That is exactly the erasure the storage rework removes elsewhere; re-introducing it in the KV API would be a step back.Scope
keyFormat/valueFormatonKvCacheStore(default-derived from the encodings, so no implementation breaks);keyFormat/valueFormatonKvCacheConfig;DefaultKvCacheStorehonours a narrow-float format (bf16/fp16 rings) as well as FP32;MemoryPlantakes the KV byte width from the store's format when one is supplied, instead ofKvCacheMode;TurboQuantKvCacheStorereports its TurboQuant format.