Skip to content

feat(memory): one view mechanism — Layout/TensorView subsumes sliced views, byte-range aliases and the packed transpose (SKEEP-003 P4, S2.1) - #1081

Merged
michalharakal merged 1 commit into
developfrom
feature/1034-one-view-mechanism
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1034-one-view-mechanism

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1034 · Phase P4 · Milestone M2 · Proposal §4.6, §4.3 rule 5

The three mechanisms

Looking at someone else's bytes had three unrelated implementations: SlicedTensorView remapped indices through an IndexMapper, BufferHandle.Aliased named a byte range, and a packed transpose rewrapped the buffer. Rule 5 says there is one: a view is a TensorView with the same Storage and a different Layout.

  • Index remap → layout arithmetic. Layout.step(axis, step) adds the strided half of the old Slice.Step (a stride multiply), so narrow + step + squeeze cover every slice kind the remapper handled. Views.kt gives DSL code the entry points: Tensor.view() / viewOrNull(), and TensorView.slice(slices) which replays the old Slice vocabulary onto those operations — returning the same view type everything else returns, so views compose instead of nesting wrappers.
  • Byte-range alias → Storage.slice. Already existed (Owner.Alias); this PR pins that it has the sharing, bounds and parent-tied semantics BufferHandle.Aliased documented, and deprecates the handle.
  • Packed transpose → a moving block axis. See below — this one turned up a bug.

The bug this found

A blocked Layout assumed the axis measured in blocks was the last one. TensorView.transpose() swapped the strides but left that assumption in place, so a transposed packed view decoded the wrong elements — zero-copy, and wrong. It was never asserted, because nothing transposed a packed view yet.

Layout now carries blockAxis, and it travels with its extent through transpose/unsqueeze/squeeze (with narrow/step consulting it). A transposed packed weight decodes correctly and still slices by whole blocks, on whichever axis carries them now.

What did not change

DefaultCpuOps.transpose keeps its O(bytes) block-grid permutation. The packed matmul kernels read a weight as input-block-major regardless of its declared shape — the bare shape swap that silently read garbage is #968/#971, and the contract is what #973 exists to write down. Until that lands, a zero-copy view cannot feed those kernels.

PackedTransposeGoldenTest pins both halves for all seven packed encodings, on JVM and Kotlin/Native: the permuted bytes are the input-block-major reordering of the original blocks, and decoded as such they carry exactly the values the zero-copy view exposes. The two paths describe the same matrix by different addressing — stated as an assertion rather than left as folklore.

Acceptance

  • OneViewMechanismTest (8 cases): every slice kind — All, Range, At, Step, multi-axis, chained — agrees with SlicedTensorView on shape and values; slice() equals the hand-composed narrow/step/squeeze; a Storage.slice alias shares memory, is bounds-checked and declares Owner.Alias; transpose and step are metadata only (same StorageId, different strides); and view() is documented as a fresh handle per call over the same array.
  • PackedTransposeGoldenTest (8 cases): shape, zero-copy (same StorageId), element-wise transposition, the block-major contract above, block-aligned narrowing of a transposed view, and seven new transpose/* goldens.
  • Golden gate green — the other 31 digests are unchanged, so no decoder or kernel moved a bit.

Deprecations

sk.ainet.lang.tensor.TensorView, SlicedTensorView and BufferHandle.Aliased are @Deprecated at WARNING level. Nothing is removed; every caller still compiles, and the files that are the old mechanism carry a file-level @Suppress("DEPRECATION") so the build stays quiet.

No ReplaceWith on the two types, deliberately: the replacement is not type-parameterized, so an IDE fix on TensorView<T, V> would produce sk.ainet.lang.memory.TensorView<T, V>, which does not compile. The message says what to write instead.

Gate

scripts/pr-gate.sh — all legs passed: jvmTest · apiCheck · verifyNpmPins jsTest wasmJsTest wasmWasiTest · linuxX64Test · assemble · :skainet-test:skainet-test-java:test.
scripts/pr-gate.sh --golden — passed on JVM and linuxX64.

No existing execution path changed — the deprecations are annotations, and the new operations are additive — so SlicingBenchmarks measures the same code it measured before.

Keeps develop green by

Deprecation over removal, and the only behavioural change is a packed-view transpose that was wrong and is now right. One API note for the dump: Layout gained a blockAxis constructor parameter, which changes its constructor signature. Layout is @ExperimentalMemoryApi, introduced in M1 and used only by the memory package, so this is a source-compatible addition with no downstream callers to break.

🤖 Generated with Claude Code

…views, byte-range aliases and the packed transpose

Closes #1034 (SKEEP-003 P4, S2.1, proposal §4.6, rule 5).

Three unrelated ways of looking at someone else's bytes existed side by
side: `SlicedTensorView` remapped indices, `BufferHandle.Aliased` named a
byte range, and a packed transpose rewrapped the buffer. All three are a
`TensorView` with a different `Layout` over the same `Storage`.

- `Layout.step(axis, step)` / `TensorView.step` — the strided half of the
  old `Slice.Step` as a stride multiply, so narrow + step + squeeze cover
  every slice kind the index remapper handled.
- `Views.kt`: `Tensor.view()` / `viewOrNull()` and
  `TensorView.slice(slices)` — the old `Slice` DSL replayed as layout
  arithmetic, returning the same view type `narrow`/`transpose`/
  `unsqueeze`/`squeeze` return, so views compose instead of nesting
  wrappers.
- `Layout` gains a **block axis**. A blocked layout used to assume its
  block-carrying axis was the last one, so transposing a packed view
  produced a view that decoded the wrong elements — the packed transpose
  was zero-copy but wrong. The block axis now travels with its extent
  through `transpose`/`unsqueeze`/`squeeze`, and `narrow`/`step` consult
  it, so a transposed packed weight decodes correctly and slices by whole
  blocks on whichever axis now carries them.
- `sk.ainet.lang.tensor.TensorView`, `SlicedTensorView` and
  `BufferHandle.Aliased` are `@Deprecated` (WARNING) with the migration
  spelled out; nothing is removed and every caller keeps compiling. No
  `ReplaceWith` on the two types: the replacement is not
  type-parameterized, so an automatic fix would emit code that does not
  compile — the message says what to write instead. Internal users of the
  old mechanism carry a file-level `@Suppress("DEPRECATION")`.

`DefaultCpuOps.transpose` still performs its O(bytes) block-grid
permutation: the packed kernels read a weight as input-block-major
whatever shape it declares (#968/#971), which is the contract #973 exists
to write down. `PackedTransposeGoldenTest` pins both halves — the
permuted bytes *are* the input-block-major reordering, and decoded as
such they carry exactly the values the zero-copy view exposes — for all
seven packed encodings, on JVM and Kotlin/Native.

`OneViewMechanismTest` compares the new mechanism against the old ones
directly: every slice kind (All/Range/At/Step, multi-axis, chained)
agrees with `SlicedTensorView` on shape and values; a `Storage.slice`
alias has the sharing, bounds and parent-tied semantics
`BufferHandle.Aliased` documented; transpose and step are metadata only;
and every derived view keeps the same `Storage`.

No existing execution path changed — the deprecations are annotations and
the new operations are additive — so `SlicingBenchmarks` measures the same
code it did before.

Gate: scripts/pr-gate.sh — all legs passed; scripts/pr-gate.sh --golden
passed on JVM and linuxX64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1081 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 53aed18 into develop Aug 24, 2026
17 checks passed
@michalharakal
michalharakal deleted the feature/1034-one-view-mechanism branch August 24, 2026 07:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S2.1] P4: one view mechanism — Layout-TensorView subsumes SlicedTensorView, BufferHandle.Aliased, packed-transpose rewrap

1 participant