Skip to content

perf(lang-core): shape-only void tracing — lazy placeholders instead of dense zeros (#1247) - #1249

Merged
michalharakal merged 1 commit into
developfrom
feat/1247-void-placeholder
Sep 2, 2026
Merged

michalharakal merged 1 commit into
developfrom
feat/1247-void-placeholder

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Part of #1247 (defect 1: VoidTensorOps allocates real dense zero buffers for every recorded op — ~9 GB of retained zeros across a 30-layer Gemma 3n E2B trace).

  • ShapeOnlyDataFactory's static-shape branch switches from DenseTensorDataFactory.zeros (which allocated twice, retained once) to placeholder — lazy zeros. Code that reads a void tensor still sees zeros (materialized on first access); the ops that are traced-and-never-read allocate nothing.
  • ShapeOnlyTensorData promoted private → public in sk.ainet.lang.tensor.data — SKaiNET-transformers hand-rolls ~9 anonymous shape-only TensorData impls that can collapse onto it (follow-up there).
  • matmulWeightTransposed override on VoidTensorOps: computes the result shape through the same calculateTransposeShape → validateMatmulShapes → calculateMatmulShape path without constructing the transposed intermediate. (Traced calls still record transpose+matmul via the KSP wrapper — the memory win is the placeholders.)

Testing: new VoidOpsAllocationTest — a weight-scale (4096×4096) op stack incl. matmulWeightTransposed asserts 0 tracked bytes under the memory tracker; read-compat test (element == 0.0f); shape parity override-vs-default. VoidTensorOpsTest (73), VoidOpsFlattenTest (21) pass unmodified; compile-dag/compile-hlo/compile-json suites green. apiDump diff is exactly the new public class.

With the sibling #1247 PRs, the full-E2B gemma3n trace+finalize completes in seconds under a 46 GB heap where it previously OOMed.

…of dense zeros (#1247)

VoidTensorOps allocated a real dense zeros buffer for every recorded op
result — twice transiently, once retained by the trace session — so a
30-layer Gemma 3n E2B trace retained ~9 GB of zeros, worst case a
weight-sized buffer per projection via matmulWeightTransposed's traced
transpose. The static-shape branch of ShapeOnlyDataFactory now returns
DenseTensorDataFactory.placeholder (LazyZero*TensorData): readers still
observe zeros, materialized and cached on first access, while a
traced-and-never-read result allocates nothing.

ShapeOnlyTensorData is promoted to public in sk.ainet.lang.tensor.data so
the hand-rolled anonymous shape-only TensorData impls in consumers
(transformer building blocks in SKaiNET-transformers) can collapse onto
it. VoidTensorOps also gains a matmulWeightTransposed override that runs
the default's exact validation and shape arithmetic without constructing
the transposed intermediate.

VoidOpsAllocationTest pins the contract under ActiveMemoryTracker: a
weight-scale projection stack tracks zero copied bytes, reads still yield
0.0f, and the override matches the default's shape semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1249 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 2d3a9d5 into develop Sep 2, 2026
19 checks passed
@michalharakal
michalharakal deleted the feat/1247-void-placeholder branch September 2, 2026 10:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant