Skip to content

feat(memory): packed TensorData façades — every GGML block format and ternary expose a zero-copy, bit-identical TensorView (SKEEP-003 P2, S1.4b) - #1069

Merged
michalharakal merged 1 commit into
developfrom
feature/1024-tensordata-facades-packed
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1024-tensordata-facades-packed

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.4b (milestone M1 #1002, PRD M1-A6): the packed encodings join the façade — every GGML block format and ternary now expose the same TensorView the dense types got in #1068, over the very bytes the loader produced.

  • PackedBlockStorage.packedView: Format(FP32, <encoding>) over a blocked Layout whose storage borrows packedData (read-only). Nothing is copied and nothing is re-ordered — that is what keeps the packed kernels bit-identical. view.get() decodes through dequantizeBlock (rule 4: never a raw byte), while the data's own get() keeps returning its code byte for source compatibility.
  • override val view on all eight packed implementations: Q4_0, Q5_0, Q5_1, Q8_0, Q4_K, Q5_K, Q6_K, Ternary2Bit.
  • Layout.blocked now also handles a block that spans rows (ternary keeps the whole tensor in one block) by addressing the flattened element sequence; a view like that refuses per-axis slicing and says why.
  • The JVM Q4/Q8MemorySegmentTensorData have no block decoder yet, so they keep the dense-only path; their view arrives with the segment kernels in [S1.7a] P3: KernelKey + rank normalization + reference matmul via defaulted KernelProvider.kernelFor(key); #993/#991 through the registry; kernel sample #1027.

Bit-identity evidence (PackedTensorDataViewTest): for every encoding the view decodes bit-identically to PackedBlockStorage.toFloatArray() — compared on raw float bits, and via a different traversal (element-wise view.get() vs block-wise dequantizeBlock), so it is a real cross-check, not a tautology. Plus: the storage is the same ByteArray instance, the view is read-only and refuses writes, whole-block slicing of a Q4_K weight matches the corresponding row of the full decode, and materialize() decodes into a Forward scope. JVM 187/187, linuxX64 73/73; the S0.9 golden parity gate (all seven GGML encodings + ternary + TurboQuant) passes unchanged.

BCV: lang-core jvm dump regenerated (additions only).

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) — all legs passed, including the golden parity tests that guard exactly this change; results in the first comment.

Closes #1024

🤖 Generated with Claude Code

… ternary expose a zero-copy, bit-identical TensorView (SKEEP-003 P2)

Milestone M1 (#1002), PRD M1-A6. Completes the façade phase started in
#1068: the packed encodings now expose the same TensorView the dense ones
do, over the very bytes the loader produced.

- PackedBlockStorage.packedView: Format(FP32, encoding) over a blocked
  Layout whose storage *borrows* packedData (read-only). Nothing is copied
  and nothing is re-ordered, which is what keeps the packed kernels
  bit-identical; view.get() decodes through dequantizeBlock (rule 4)
  while the data's own get() keeps returning its code byte for source
  compatibility.
- override val view on all eight packed implementations: Q4_0, Q5_0,
  Q5_1, Q8_0, Q4_K, Q5_K, Q6_K and Ternary2Bit.
- Layout.blocked handles a block that spans rows (ternary keeps the whole
  tensor in one block) by addressing the flattened element sequence;
  TensorView refuses to slice such a view per axis and says why.
  (The JVM Q4/Q8 MemorySegment data have no block decoder yet — they get
  their view with the segment kernels in #1027.)
- PackedTensorDataViewTest: for every encoding the view decodes
  bit-identically to PackedBlockStorage.toFloatArray() (raw-bit
  comparison, element-wise vs block-wise traversal), borrows the same
  byte array, is read-only, refuses writes; whole-block slicing of a Q4_K
  weight matches the corresponding row; materialize() decodes into a
  Forward scope. JVM 187/187, linuxX64 73/73. BCV dumps regenerated.

Closes #1024

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on d4b1ef3: all legs passed — jvmTest (incl. the golden parity tests) · apiCheck · JS/Wasm 2m44s · linuxX64Test 1m22s · assemble · Java API tests.

Targeted: JVM 187/187 and linuxX64 73/73 (sk.ainet.lang.memory.* + sk.ainet.lang.tensor.data.*), including the per-encoding raw-bit decode comparison.

@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1069 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit ed522de into develop Aug 23, 2026
17 checks passed
@michalharakal
michalharakal deleted the feature/1024-tensordata-facades-packed branch August 23, 2026 16:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.4b] P2: packed TensorData implementations (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary, TurboQuant) become façades over TensorView

1 participant