Skip to content

feat(memory): Format(dtype, encoding), TensorData.encoding, Tensor/TensorStorage.format (SKEEP-003 P1, S0.4) - #1051

Merged
michalharakal merged 3 commits into
developfrom
feature/1008-format-and-encoding
Aug 23, 2026
Merged

michalharakal merged 3 commits into
developfrom
feature/1008-format-and-encoding

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S0.4 (milestone M0 #1001, acceptance M0-A3): every tensor reports a coherent Format(dtype, encoding) — a packed weight is logically FP32 with its block encoding, never a Byte-typed tensor.

  • New package sk.ainet.lang.memory (opt-in @ExperimentalMemoryApi while M0–M1 shape the API; this is where AllocationSpec, Storage, Scope, TensorView … will live): data class Format(dtype: DType, encoding: TensorEncoding) with isDense, physicalBytes(count), Format.dense(dtype), toString() = Float32/Q4_K.
  • TensorData.encoding: TensorEncoding? as a default member (null = plain dense representation of the dtype). TensorData is implemented outside lang-core, so no abstract member is added. The packed implementations already expose encoding via PackedBlockStorage; narrow-float data now reports Dense(2) (physical width regardless of the dtype witness); the JVM Q4/Q8MemorySegmentTensorData report Q4_0/Q8_0.
  • Tensor<*, *>.format / formatOrNull = Format(DType.fromWitness(dtype), data.encoding ?: Dense(dtype.sizeInBytes)) (uses [S0.2a] P0: DType carries its witness (DType.witness, fromWitness, entries); descriptors/readers report DType additively #1007's fromWitness); TensorStorage.format.
  • FormatTest (commonTest): dense and narrow data, all seven GGML packed encodings + ternary (Format(FP32, <enc>), string form), storage format, non-concrete witness.
  • BCV: lang-core jvm dump regenerated (additions only).

Stacked on #1050 (S0.2a — DType.fromWitness); base is that branch and GitHub retargets to develop when it merges.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25) on the stacked tree; results in the first comment. Targeted: lang-core sk.ainet.lang.memory.* + sk.ainet.lang.tensor.data.* 110/110.

Closes #1008

🤖 Generated with Claude Code

…mber; Tensor/TensorStorage.format (SKEEP-003 P1)

Milestone M0 (#1001), acceptance M0-A3: every tensor reports a coherent
Format — a packed weight is logically FP32 with its block encoding, never
a Byte-typed tensor.

- New package sk.ainet.lang.memory (opt-in @ExperimentalMemoryApi while
  M0–M1 shape the API): data class Format(dtype: DType, encoding:
  TensorEncoding) with isDense, physicalBytes(count), Format.dense(dtype),
  toString "Float32/Q4_K".
- TensorData.encoding: TensorEncoding? as a *default* member (null =
  plain dense representation of the dtype) — TensorData is implemented
  outside lang-core, so no abstract member. The packed implementations
  already expose `encoding` through PackedBlockStorage; NarrowFloat data
  now reports Dense(2) (physical width whatever the dtype witness), the
  JVM Q4/Q8 MemorySegment data report Q4_0/Q8_0.
- Tensor<*, *>.format / formatOrNull: Format(DType.fromWitness(dtype),
  data.encoding ?: Dense(dtype.sizeInBytes)); TensorStorage.format.
- FormatTest (commonTest): dense/narrow/packed data for all seven GGML
  encodings + ternary, storage format, non-concrete witness.
- BCV: lang-core jvm dump regenerated (additions only).

Closes #1008

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on e8a0efe (stacked on #1050):

=== pr-gate: JVM tests ===                                      BUILD SUCCESSFUL in 4m 18s
=== pr-gate: apiCheck ===                                       BUILD SUCCESSFUL in 1m 15s
=== pr-gate: verifyNpmPins jsTest wasmJsTest wasmWasiTest ===   BUILD SUCCESSFUL in 2m 39s
=== pr-gate: linuxX64Test ===                                   BUILD SUCCESSFUL in 1m 26s
=== pr-gate: assemble ===                                       BUILD SUCCESSFUL in 2m 12s
=== pr-gate: :skainet-test:skainet-test-java:test ===           BUILD SUCCESSFUL in 21s
pr-gate: all legs passed.

Targeted: lang-core sk.ainet.lang.memory.* + sk.ainet.lang.tensor.data.* 110/110.

@michalharakal
michalharakal changed the base branch from feature/1007-dtype-witness to develop August 22, 2026 19:18
…torageSpec (SKEEP-003 P0)

Milestone M0 (#1001), SKEEP-003 prerequisite "StorageSpec becomes the
allocation spec".

- sk.ainet.lang.memory.AllocationSpec(format, elementCount, domain =
  HOST_HEAP, scope = AMBIENT, mutable = true, alignment = 64): bytes /
  bytesOrNull from the encoding, AllocationSpec.of(format, shape, ...),
  validation (count >= 0, power-of-two alignment). The input of the M0
  MemoryPlan line items and, from M1, of Storage.allocate(spec, scope).
- enum ScopeKind { MODEL, FORWARD, AMBIENT } — the lifetime classes of
  SKEEP-003 §4.5, declared now; M1's Scope exposes kind: ScopeKind.
- StorageSpec and its companion factories @deprecated(ReplaceWith
  AllocationSpec); zero consumers outside its own file and the S0.1 bridge
  test (kept on purpose, suppressed). StorageSpec.toAllocationSpec(count)
  is the migration path. Removed at the next major.
- AllocationSpecTest; BCV lang-core jvm dump regenerated (additions +
  deprecation annotations).

Closes #1009

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ion-spec

feat(memory): AllocationSpec + ScopeKind replace the never-consumed StorageSpec (SKEEP-003 P0, S0.3)
@michalharakal
michalharakal requested a review from aharakal August 22, 2026 20:16
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1051 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S0.4] P1: Format(dtype, encoding) + TensorData.encoding default member — every tensor reports its Format (Q4_K → F32/Q4_K)

2 participants