feat(memory): packed TensorData façades — every GGML block format and ternary expose a zero-copy, bit-identical TensorView (SKEEP-003 P2, S1.4b) - #1069
Merged
Conversation
… ternary expose a zero-copy, bit-identical TensorView (SKEEP-003 P2) Milestone M1 (#1002), PRD M1-A6. Completes the façade phase started in #1068: the packed encodings now expose the same TensorView the dense ones do, over the very bytes the loader produced. - PackedBlockStorage.packedView: Format(FP32, encoding) over a blocked Layout whose storage *borrows* packedData (read-only). Nothing is copied and nothing is re-ordered, which is what keeps the packed kernels bit-identical; view.get() decodes through dequantizeBlock (rule 4) while the data's own get() keeps returning its code byte for source compatibility. - override val view on all eight packed implementations: Q4_0, Q5_0, Q5_1, Q8_0, Q4_K, Q5_K, Q6_K and Ternary2Bit. - Layout.blocked handles a block that spans rows (ternary keeps the whole tensor in one block) by addressing the flattened element sequence; TensorView refuses to slice such a view per axis and says why. (The JVM Q4/Q8 MemorySegment data have no block decoder yet — they get their view with the segment kernels in #1027.) - PackedTensorDataViewTest: for every encoding the view decodes bit-identically to PackedBlockStorage.toFloatArray() (raw-bit comparison, element-wise vs block-wise traversal), borrows the same byte array, is read-only, refuses writes; whole-block slicing of a Q4_K weight matches the corresponding row; materialize() decodes into a Forward scope. JVM 187/187, linuxX64 73/73. BCV dumps regenerated. Closes #1024 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Local gate Targeted: JVM 187/187 and linuxX64 73/73 ( |
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S1.4b (milestone M1 #1002, PRD M1-A6): the packed encodings join the façade — every GGML block format and ternary now expose the same
TensorViewthe dense types got in #1068, over the very bytes the loader produced.PackedBlockStorage.packedView:Format(FP32, <encoding>)over a blockedLayoutwhose storage borrowspackedData(read-only). Nothing is copied and nothing is re-ordered — that is what keeps the packed kernels bit-identical.view.get()decodes throughdequantizeBlock(rule 4: never a raw byte), while the data's ownget()keeps returning its code byte for source compatibility.override val viewon all eight packed implementations:Q4_0,Q5_0,Q5_1,Q8_0,Q4_K,Q5_K,Q6_K,Ternary2Bit.Layout.blockednow also handles a block that spans rows (ternary keeps the whole tensor in one block) by addressing the flattened element sequence; a view like that refuses per-axis slicing and says why.Q4/Q8MemorySegmentTensorDatahave no block decoder yet, so they keep the dense-only path; their view arrives with the segment kernels in [S1.7a] P3:KernelKey+ rank normalization + reference matmul via defaultedKernelProvider.kernelFor(key); #993/#991 through the registry; kernel sample #1027.Bit-identity evidence (
PackedTensorDataViewTest): for every encoding the view decodes bit-identically toPackedBlockStorage.toFloatArray()— compared on raw float bits, and via a different traversal (element-wiseview.get()vs block-wisedequantizeBlock), so it is a real cross-check, not a tautology. Plus: the storage is the sameByteArrayinstance, the view is read-only and refuses writes, whole-block slicing of a Q4_K weight matches the corresponding row of the full decode, andmaterialize()decodes into aForwardscope. JVM 187/187, linuxX64 73/73; the S0.9 golden parity gate (all seven GGML encodings + ternary + TurboQuant) passes unchanged.BCV: lang-core jvm dump regenerated (additions only).
Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25) — all legs passed, including the golden parity tests that guard exactly this change; results in the first comment.Closes #1024
🤖 Generated with Claude Code