Skip to content

feat(memory): TensorData exposes a zero-copy TensorView — dense, narrow-float and MemorySegment façades (SKEEP-003 P2, S1.4a) - #1068

Merged
michalharakal merged 1 commit into
developfrom
feature/1023-tensordata-facades-dense
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1023-tensordata-facades-dense

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.4a (milestone M1 #1002, PRD M1-A6 / M1-A9): every dense TensorData implementation becomes a façade over TensorView — additively, with no behaviour change and no hot-path change.

  • TensorData.view: TensorView? default member (null when an implementation cannot expose one). The view is over the same bytes: array-backed data borrows its array via Storage.Heap.wrap, so writes through either side are visible on both and nothing is copied.
  • Design note (Phase-2 spike bench(spike): Phase-2 TensorView/Storage access-path spike — JMH benchmarks, flat-RSS loop, report (SKEEP-003 decision #6, S1.0) #1060): get/set are not routed through the view. The spike measured per-element access through a view at +80 % (heap) — a view is for unwrapping once per call, not for per-element reads. So TensorData keeps its own fast path and exposes the view for migrated kernels; the reference matmul/elementwise paths are byte-for-byte the same code as before.
  • Providers: FloatArrayTensorData / IntArrayTensorData (dense FP32 / Int32 views — covers DenseFloatArray, DenseInt, LazyZero*), NarrowFloatDenseTensorData (a Dense(2) view over the packed bytes decoded by the data's own codec, via the new NarrowFloatDecoder), MemorySegmentTensorData (SegmentStorage.borrow over the same segment, sliced at segmentByteOffset).
  • TensorView.get() now prefers a decoder when one is present, so a narrow-float view (physically Dense(2), logically FP16/BF16) decodes instead of taking the plain float path.
  • LazyMaterializationStrategy's private view field renamed to sourceView — it shadowed the new member (the old sk.ainet.lang.tensor.TensorView and the new sk.ainet.lang.memory.TensorView coexist until [S2.1] P4: one view mechanism — Layout-TensorView subsumes SlicedTensorView, BufferHandle.Aliased, packed-transpose rewrap #1034).
  • TensorDataViewTest: zero-copy identity of the borrowed array, reads/writes agreeing in both directions, lazy-zero materialization, BF16/FP16 decode parity with the data's own get(), null for data without a view. JVM 177/177, linuxX64 63/63.
  • BCV: lang-core jvm dump regenerated (additions only).

Packed façades (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary) are #1024 and are gated by the golden parity tests staying bit-identical.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) — all functional legs passed: jvmTest · apiCheck · JS/Wasm · linuxX64Test · assemble · Java API tests. A targeted JMH run (elementwise / quantized matmul / reductions) against docs/design/memory/baseline-2026-08-22.md follows in a comment as hot-path evidence.

Closes #1023

🤖 Generated with Claude Code

…ow-float and MemorySegment façades (SKEEP-003 P2)

Milestone M1 (#1002), PRD M1-A6/M1-A9. SKEEP-003 §4.1: every TensorData
implementation becomes a façade over TensorView, additively and without a
behaviour change.

- TensorData.view: TensorView? default member (null when an implementation
  cannot expose one). The view is over the *same* bytes — array-backed data
  borrows its array via Storage.Heap.wrap — so writes through either side
  are visible on both and nothing is copied. Per-element access stays on
  TensorData's own fast path: the Phase-2 spike (#1016) showed a view is
  for unwrapping once per call, not for per-element reads, so no hot path
  changes and the benchmarks are untouched.
- Providers: FloatArrayTensorData / IntArrayTensorData (dense FP32 / Int32
  views, covering DenseFloatArray/DenseInt/LazyZero*), NarrowFloatDense
  (Dense(2) view over the packed bytes decoded by the data's codec, new
  NarrowFloatDecoder), MemorySegmentTensorData (SegmentStorage.borrow over
  the same segment, sliced at segmentByteOffset).
- TensorView.get() prefers a decoder when present, so a narrow-float view
  (Dense(2) but 16-bit encoded) decodes instead of taking the plain path.
- LazyMaterializationStrategy's private `view` field renamed to
  `sourceView`: it shadowed the new TensorData.view member.
- TensorDataViewTest: zero-copy identity of the borrowed array, reads and
  writes agreeing both ways, lazy-zero materialization, BF16/FP16 decode
  parity with the data's own get(), and null for data that has no view.
  JVM 177/177, linuxX64 63/63. BCV: lang-core jvm dump regenerated.

Closes #1023

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on 68aa852:

=== pr-gate: JVM tests ===                                      BUILD SUCCESSFUL in 7m 28s
=== pr-gate: apiCheck ===                                       BUILD SUCCESSFUL in 18s
=== pr-gate: verifyNpmPins jsTest wasmJsTest wasmWasiTest ===   BUILD SUCCESSFUL
=== pr-gate: linuxX64Test ===                                   BUILD SUCCESSFUL
=== pr-gate: assemble ===                                       BUILD SUCCESSFUL in 2m 6s
=== pr-gate: :skainet-test:skainet-test-java:test ===           BUILD SUCCESSFUL in 21s

Targeted: JVM 177/177 and linuxX64 63/63 (sk.ainet.lang.memory.* + sk.ainet.lang.tensor.data.*). JMH hot-path numbers follow.

@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1068 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal

Copy link
Copy Markdown
Contributor Author

Hot-path evidence — targeted JMH on this branch (68aa852) vs the committed baseline docs/design/memory/baseline-2026-08-22.md (i7-9750H, JDK 25, fork 1 / 3 warmup / 5 iterations, machine idle):

Benchmark Params Mode Pre-M0 Post-M0 Δ Units
add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=false thrpt 1.217 ± 0.011 1.226 ± 0.002 -0.7% ops/ms
add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=true thrpt 1.705 ± 0.030 1.727 ± 0.035 -1.3% ops/ms
matmul_q4k_panama QuantizedMatmulBench shape=1024-1024 avgt 0.112 ± 0.017 0.110 ± 0.021 -1.8% ms/op
matmul_q4k_panama QuantizedMatmulBench shape=4096-1024 avgt 0.336 ± 0.003 0.340 ± 0.014 +1.2% ms/op
matmul_q4k_panama QuantizedMatmulBench shape=4096-4096 avgt 1.223 ± 0.017 1.249 ± 0.026 +2.1% ms/op
mean_1M_fp32 Reductions1MBench vectorEnabled=false thrpt 1.026 ± 0.004 1.026 ± 0.006 +0.0% ops/ms
mean_1M_fp32 Reductions1MBench vectorEnabled=true thrpt 1.026 ± 0.011 1.027 ± 0.007 -0.1% ops/ms
sum_1M_fp32 Reductions1MBench vectorEnabled=false thrpt 1.027 ± 0.006 1.026 ± 0.008 +0.1% ops/ms
sum_1M_fp32 Reductions1MBench vectorEnabled=true thrpt 1.026 ± 0.003 1.023 ± 0.004 +0.3% ops/ms

worst slowdown: +2.1% (positive = slower after M0)

Everything is within ±2.1 %, inside the error bars — expected, since this slice adds a view property and changes no get/set code (the Phase-2 spike's rule: unwrap once per call, never per element). Decision #6's elementwise ≤ 3 % budget holds: add_1M_fp32 is −1.3 % (i.e. slightly faster) with the vector path and −0.7 % without.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.4a] P2: dense TensorData implementations become façades over TensorView (bit-identical)

1 participant