perf(lang-core): shape-only void tracing — lazy placeholders instead of dense zeros (#1247) - #1249
Merged
Merged
Conversation
…of dense zeros (#1247) VoidTensorOps allocated a real dense zeros buffer for every recorded op result — twice transiently, once retained by the trace session — so a 30-layer Gemma 3n E2B trace retained ~9 GB of zeros, worst case a weight-sized buffer per projection via matmulWeightTransposed's traced transpose. The static-shape branch of ShapeOnlyDataFactory now returns DenseTensorDataFactory.placeholder (LazyZero*TensorData): readers still observe zeros, materialized and cached on first access, while a traced-and-never-read result allocates nothing. ShapeOnlyTensorData is promoted to public in sk.ainet.lang.tensor.data so the hand-rolled anonymous shape-only TensorData impls in consumers (transformer building blocks in SKaiNET-transformers) can collapse onto it. VoidTensorOps also gains a matmulWeightTransposed override that runs the default's exact validation and shape arithmetic without constructing the transposed intermediate. VoidOpsAllocationTest pins the contract under ActiveMemoryTracker: a weight-scale projection stack tracks zero copied bytes, reads still yield 0.0f, and the override matches the default's shape semantics. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This was referenced Sep 2, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #1247 (defect 1:
VoidTensorOpsallocates real dense zero buffers for every recorded op — ~9 GB of retained zeros across a 30-layer Gemma 3n E2B trace).ShapeOnlyDataFactory's static-shape branch switches fromDenseTensorDataFactory.zeros(which allocated twice, retained once) toplaceholder— lazy zeros. Code that reads a void tensor still sees zeros (materialized on first access); the ops that are traced-and-never-read allocate nothing.ShapeOnlyTensorDatapromoted private → public insk.ainet.lang.tensor.data— SKaiNET-transformers hand-rolls ~9 anonymous shape-onlyTensorDataimpls that can collapse onto it (follow-up there).matmulWeightTransposedoverride onVoidTensorOps: computes the result shape through the samecalculateTransposeShape→validateMatmulShapes→calculateMatmulShapepath without constructing the transposed intermediate. (Traced calls still record transpose+matmul via the KSP wrapper — the memory win is the placeholders.)Testing: new
VoidOpsAllocationTest— a weight-scale (4096×4096) op stack incl.matmulWeightTransposedasserts 0 tracked bytes under the memory tracker; read-compat test (element == 0.0f); shape parity override-vs-default.VoidTensorOpsTest(73),VoidOpsFlattenTest(21) pass unmodified; compile-dag/compile-hlo/compile-json suites green. apiDump diff is exactly the new public class.With the sibling #1247 PRs, the full-E2B gemma3n trace+finalize completes in seconds under a 46 GB heap where it previously OOMed.