Eager op outputs draw from the active scope - #1168
Merged
Merged
Conversation
…kt's deletion #1142 deleted MemoryPlannerTest.kt for the dead planner it tested — but the file also housed MemoryTrackerTest, whose four unit tests (trackAndReport, trackCopies, clearResetsState, fileBackedTracking) cover the live MemoryTracker aggregate reports. Restored verbatim in its own file; only the planner-owned imports are dropped. Part of #1146's memory-accounting groundwork. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The creation path learned to consult ExecutionContext.memoryScope in #1145; op outputs still reached the GC because DefaultCpuOps allocates them through its TensorDataFactory. This closes that half (#1146): - TensorDataFactory.adoptFloatArray — the op-output entry point. Unlike wrapFloatArray, the buffer is a fresh, ownership-transferred result, which licenses a factory to relocate it. DenseTensorDataFactory adopts FP32/FP16 zero-copy (one copy FEWER than before on the common path, which paid fromFloatArray's copyOf); narrow dtypes keep their tagged fromFloatArray path. - ScopedTensorDataFactory decorates any factory: under a non-Ambient scope, dense-FP32 zeros/ones/full/init/fromFloatArray/adoptFloatArray land in the slab as StorageFloatTensorData; the region is always fully written because a reset slab is dirty. wrap* stays un-intercepted — its zero-copy caller-owned contract must not die at reset(). - ExecutionContext.withTensorDataFactory + the DirectCpuExecutionContext override let ScopedExecutionContext rebuild the base around the scoped factory and re-bind created tensors to the rebuilt ops, so 'a + b' dispatches into scope-allocating ops. - The FP32 fast paths are offset-aware (floatWindowOf): slab-backed operands stay on the flat primitive loops instead of falling to the boxed generic path (#949's 83%-overhead cliff — live since #1145 for scope-created tensors, now closed). Tail sites read slab data as its exact logical window through floatBufferOf (one copy, still far off the boxed path). DefaultCpuOpsJvm's 18 direct constructions route through floatResult; its Panama vector paths fall back to the common window loops for slab operands (follow-up: offset-aware vectors). - KernelDispatch.matmul receives dispatchScope(), so requantize/prepack adapter allocations land in the slab too. - ScopedOpOutputsTest pins: slab-backed outputs, flat peakFloats across 8 steps with zero overflow (steady-state decode allocates no new slab bytes), bit-identical numerics vs Ambient, loud StorageClosedException on stale reads, and the overflow path. Also restores MemoryTrackerTest (lost with MemoryPlannerTest.kt in #1142). Closes #1146. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1146 (parent #1135).
Op outputs draw from the active scope — the other half of the Scope split (#1145 wired creation).
The seam:
TensorDataFactory.adoptFloatArray, a new op-output entry point. UnlikewrapFloatArray, the buffer is a fresh, ownership-transferred result — which licenses a factory to relocate it.DenseTensorDataFactoryadopts FP32/FP16 zero-copy (one copy fewer than before: the common path previously paidfromFloatArray'scopyOf); narrow dtypes keep their tagged path.ScopedTensorDataFactorydecorates any factory: under a non-Ambient scope, dense-FP32zeros/ones/full/init/fromFloatArray/adoptFloatArrayland in the slab (always fully written — a reset slab is dirty).wrap*stays un-intercepted: its zero-copy caller-owned contract must not die atreset().The wiring:
ExecutionContext.withTensorDataFactory(overridden byDirectCpuExecutionContext) letsScopedExecutionContextrebuild its base around the scoped factory, and itsfromDataoverride re-binds created tensors to the rebuilt ops — soa + bdispatches into scope-allocating ops. Tensors made before entering the scope keep their unscoped ops, deliberately: their results must not die atreset().KernelDispatch.matmulreceivesdispatchScope(), so requantize/prepack adapter allocations land in the slab too.The cliff guard: the FP32 fast paths are offset-aware (
floatWindowOf) — slab-backed operands stay on the flat primitive loops instead of falling to the boxed generic path (#949's 83 %-overhead cliff, live since #1145 for scope-created tensors, now closed). Tail sites read slab data as its exact logical window throughfloatBufferOf(one copy, still far off the boxed path).DefaultCpuOpsJvm's 18 directDenseFloatArrayTensorDataconstructions route through thefloatResultfunnel; two deliberate keepers (the 1-float reduce scalar, andmean's in-place divide on it) stay direct.Proof (
ScopedOpOutputsTest): slab-backed outputs under a scope and plain-array outputs under Ambient; an 8-step decode loop with flatpeakFloatsand zero overflow — steady-state eager decode allocates no new slab bytes; bit-identical numerics vs Ambient (including the slab-overflow path); loudStorageClosedExceptionon stale reads. Note: the existingDecodeHarnessnever touchesTensorOps(it is a pure View/KernelDispatch harness that already usedForwardScope), so this test is the eager-ops equivalent of its M1-A1/M1-A3 assertions rather than an extension of it.Also restores
MemoryTrackerTest, lost as a co-tenant ofMemoryPlannerTest.kt's deletion in #1142.Follow-up filed as a note here: offset-aware Panama vector loops —
DefaultCpuOpsJvm's vector paths return null for slab operands and fall back to the common scalar-window loops (correct, slower);FloatVector.fromArray(species, arr, off)supports offsets, so the upgrade is mechanical.Full pr-gate green.
🤖 Generated with Claude Code