feat(memory): Layout + TensorView — zero-copy views over Storage, decoding get(), materialize() as the only copy point (SKEEP-003 P2, S1.3) - #1066
Merged
Conversation
…oding get(), materialize() as the only copy point (SKEEP-003 P2) Milestone M1 (#1002), PRD M1-F6. SKEEP-003 §4.2, rules 4–6: the view is the only thing a kernel receives, it never owns bytes, and every slice / transpose / unsqueeze is another view over the same Storage. - sk.ainet.lang.memory.Layout: strides (in elements, or in *blocks* for a packed format), offsetElements/offsetBytes, isRowMajor, isContiguous (unit extents ignored), indexOf/byteOffsetOf, and the metadata-only narrow / transpose / unsqueeze / squeeze; Layout.rowMajor(shape, format) and Layout.blocked(shape, blockSize, bytesPerBlock) — the latter is what makes a packed weight sliceable and transposable without touching bytes. - sk.ainet.lang.memory.TensorView(shape, format, layout, storage, id): narrow/transpose/unsqueeze/squeeze return views over the same storage (ids derive as `w[1..3)`), get() returns the decoded logical value for packed encodings and never a raw byte (rule 4), set() is refused on a packed or read-only view, toFloatArray() is the reference materialization, materialize(format, scope) is the single copy point (rule 6) and allocates in the target scope. Narrowing the block axis of a packed view must align to whole blocks. - BlockDecoder + PackedBlockDecoder: the bridge from today's PackedBlockStorage implementations (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary) to TensorView; M2's Encoding descriptors will implement the same interface. - LayoutTest and TensorViewTest (dense read/write, shared-storage views, packed decode vs PackedBlockStorage.toFloatArray, whole-block slicing, materialize as a copy, views over closed storage). JVM 67/67, linuxX64 58/58 memory tests. BCV: lang-core jvm dump regenerated. Closes #1022 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Local gate Targeted: JVM 67/67 and linuxX64 58/58 |
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S1.3 (milestone M1 #1002, PRD M1-F6):
Layout+TensorView— the type a kernel actually receives (§4.2, rules 4–6). A view never owns bytes; every slice / transpose / unsqueeze is another view over the sameStorage;materialize()is the only copy point.Layout: strides in elements — or in blocks for a packed format — plusoffsetElements/offsetBytes,isRowMajor,isContiguous(unit extents ignored),indexOf/byteOffsetOf, and the metadata-onlynarrow/transpose/unsqueeze/squeeze.Layout.rowMajor(shape, format)andLayout.blocked(shape, blockSize, bytesPerBlock)— the latter is what lets a packed weight be sliced and transposed without touching bytes (rule 5; the packed-transpose trick becomes "a layout that says transposed over an encoding that says blocked").TensorView(shape, format, layout, storage, id):narrow/transpose/unsqueeze/squeezereturn views over the same storage (ids derive asmodel.w[1..3)),get()returns the decoded logical value for packed encodings and never a raw byte (rule 4 — the Quantized matmul dispatch skips rank-1 (single-token decode) activations, falls through to broken matmulGeneric #993/Quantized matmul dispatch silently falls back to a broken generic kernel for rank>2 attention projections (ClassCastException: Byte cannot be cast to Float) #991 class becomes impossible by construction),set()is refused on packed or read-only views,toFloatArray()is the reference materialization, andmaterialize(format, scope)is the single copy point (rule 6), allocating in the targetScope. Narrowing the block axis of a packed view must align to whole blocks.BlockDecoder+PackedBlockDecoder: the bridge from today'sPackedBlockStorageimplementations (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary) into views; M2'sEncodingdescriptors will implement the same interface.LayoutTest(strides/offsets/contiguity, narrow keeps strides, transpose is an involution, unsqueeze↔squeeze, blocked geometry) andTensorViewTest(dense read/write, views sharing storage and bytes, a Q8_0 view decoding identically toPackedBlockStorage.toFloatArray(), whole-block slicing of a Q4_K weight, materialize as a real copy incl. packed→dense, views over closed storage refused). JVM 67/67, linuxX64 58/58sk.ainet.lang.memory.*tests.TensorDatafaçades are [S1.4a] P2: denseTensorDataimplementations become façades overTensorView(bit-identical) #1023/[S1.4b] P2: packedTensorDataimplementations (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary, TurboQuant) become façades overTensorView#1024.Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25) — all legs passed; results in the first comment.Closes #1022
🤖 Generated with Claude Code