feat(memory): TensorData exposes a zero-copy TensorView — dense, narrow-float and MemorySegment façades (SKEEP-003 P2, S1.4a) - #1068
Conversation
…ow-float and MemorySegment façades (SKEEP-003 P2) Milestone M1 (#1002), PRD M1-A6/M1-A9. SKEEP-003 §4.1: every TensorData implementation becomes a façade over TensorView, additively and without a behaviour change. - TensorData.view: TensorView? default member (null when an implementation cannot expose one). The view is over the *same* bytes — array-backed data borrows its array via Storage.Heap.wrap — so writes through either side are visible on both and nothing is copied. Per-element access stays on TensorData's own fast path: the Phase-2 spike (#1016) showed a view is for unwrapping once per call, not for per-element reads, so no hot path changes and the benchmarks are untouched. - Providers: FloatArrayTensorData / IntArrayTensorData (dense FP32 / Int32 views, covering DenseFloatArray/DenseInt/LazyZero*), NarrowFloatDense (Dense(2) view over the packed bytes decoded by the data's codec, new NarrowFloatDecoder), MemorySegmentTensorData (SegmentStorage.borrow over the same segment, sliced at segmentByteOffset). - TensorView.get() prefers a decoder when present, so a narrow-float view (Dense(2) but 16-bit encoded) decodes instead of taking the plain path. - LazyMaterializationStrategy's private `view` field renamed to `sourceView`: it shadowed the new TensorData.view member. - TensorDataViewTest: zero-copy identity of the borrowed array, reads and writes agreeing both ways, lazy-zero materialization, BF16/FP16 decode parity with the data's own get(), and null for data that has no view. JVM 177/177, linuxX64 63/63. BCV: lang-core jvm dump regenerated. Closes #1023 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Local gate Targeted: JVM 177/177 and linuxX64 63/63 ( |
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
|
Hot-path evidence — targeted JMH on this branch (68aa852) vs the committed baseline
worst slowdown: +2.1% (positive = slower after M0) Everything is within ±2.1 %, inside the error bars — expected, since this slice adds a |
Summary
SKEEP-003 slice S1.4a (milestone M1 #1002, PRD M1-A6 / M1-A9): every dense
TensorDataimplementation becomes a façade overTensorView— additively, with no behaviour change and no hot-path change.TensorData.view: TensorView?default member (nullwhen an implementation cannot expose one). The view is over the same bytes: array-backed data borrows its array viaStorage.Heap.wrap, so writes through either side are visible on both and nothing is copied.get/setare not routed through the view. The spike measured per-element access through a view at +80 % (heap) — a view is for unwrapping once per call, not for per-element reads. SoTensorDatakeeps its own fast path and exposes the view for migrated kernels; the reference matmul/elementwise paths are byte-for-byte the same code as before.FloatArrayTensorData/IntArrayTensorData(dense FP32 / Int32 views — coversDenseFloatArray,DenseInt,LazyZero*),NarrowFloatDenseTensorData(aDense(2)view over the packed bytes decoded by the data's own codec, via the newNarrowFloatDecoder),MemorySegmentTensorData(SegmentStorage.borrowover the same segment, sliced atsegmentByteOffset).TensorView.get()now prefers a decoder when one is present, so a narrow-float view (physicallyDense(2), logically FP16/BF16) decodes instead of taking the plain float path.LazyMaterializationStrategy's privateviewfield renamed tosourceView— it shadowed the new member (the oldsk.ainet.lang.tensor.TensorViewand the newsk.ainet.lang.memory.TensorViewcoexist until [S2.1] P4: one view mechanism — Layout-TensorViewsubsumesSlicedTensorView,BufferHandle.Aliased, packed-transpose rewrap #1034).TensorDataViewTest: zero-copy identity of the borrowed array, reads/writes agreeing in both directions, lazy-zero materialization, BF16/FP16 decode parity with the data's ownget(),nullfor data without a view. JVM 177/177, linuxX64 63/63.Packed façades (Q4_0…Q8_0, Q4_K/Q5_K/Q6_K, ternary) are #1024 and are gated by the golden parity tests staying bit-identical.
Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25) — all functional legs passed: jvmTest · apiCheck · JS/Wasm · linuxX64Test · assemble · Java API tests. A targeted JMH run (elementwise / quantized matmul / reductions) againstdocs/design/memory/baseline-2026-08-22.mdfollows in a comment as hot-path evidence.Closes #1023
🤖 Generated with Claude Code