Skip to content

Eager op outputs draw from the active scope - #1168

Merged
michalharakal merged 2 commits into
developfrom
feature/1146-op-outputs-through-scope
Aug 26, 2026
Merged

michalharakal merged 2 commits into
developfrom
feature/1146-op-outputs-through-scope

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1146 (parent #1135).

Op outputs draw from the active scope — the other half of the Scope split (#1145 wired creation).

The seam: TensorDataFactory.adoptFloatArray, a new op-output entry point. Unlike wrapFloatArray, the buffer is a fresh, ownership-transferred result — which licenses a factory to relocate it. DenseTensorDataFactory adopts FP32/FP16 zero-copy (one copy fewer than before: the common path previously paid fromFloatArray's copyOf); narrow dtypes keep their tagged path. ScopedTensorDataFactory decorates any factory: under a non-Ambient scope, dense-FP32 zeros/ones/full/init/fromFloatArray/adoptFloatArray land in the slab (always fully written — a reset slab is dirty). wrap* stays un-intercepted: its zero-copy caller-owned contract must not die at reset().

The wiring: ExecutionContext.withTensorDataFactory (overridden by DirectCpuExecutionContext) lets ScopedExecutionContext rebuild its base around the scoped factory, and its fromData override re-binds created tensors to the rebuilt ops — so a + b dispatches into scope-allocating ops. Tensors made before entering the scope keep their unscoped ops, deliberately: their results must not die at reset(). KernelDispatch.matmul receives dispatchScope(), so requantize/prepack adapter allocations land in the slab too.

The cliff guard: the FP32 fast paths are offset-aware (floatWindowOf) — slab-backed operands stay on the flat primitive loops instead of falling to the boxed generic path (#949's 83 %-overhead cliff, live since #1145 for scope-created tensors, now closed). Tail sites read slab data as its exact logical window through floatBufferOf (one copy, still far off the boxed path). DefaultCpuOpsJvm's 18 direct DenseFloatArrayTensorData constructions route through the floatResult funnel; two deliberate keepers (the 1-float reduce scalar, and mean's in-place divide on it) stay direct.

Proof (ScopedOpOutputsTest): slab-backed outputs under a scope and plain-array outputs under Ambient; an 8-step decode loop with flat peakFloats and zero overflow — steady-state eager decode allocates no new slab bytes; bit-identical numerics vs Ambient (including the slab-overflow path); loud StorageClosedException on stale reads. Note: the existing DecodeHarness never touches TensorOps (it is a pure View/KernelDispatch harness that already used ForwardScope), so this test is the eager-ops equivalent of its M1-A1/M1-A3 assertions rather than an extension of it.

Also restores MemoryTrackerTest, lost as a co-tenant of MemoryPlannerTest.kt's deletion in #1142.

Follow-up filed as a note here: offset-aware Panama vector loops — DefaultCpuOpsJvm's vector paths return null for slab operands and fall back to the common scalar-window loops (correct, slower); FloatVector.fromArray(species, arr, off) supports offsets, so the upgrade is mechanical.

Full pr-gate green.

🤖 Generated with Claude Code

michalharakal and others added 2 commits August 26, 2026 14:24
…kt's deletion

#1142 deleted MemoryPlannerTest.kt for the dead planner it tested — but
the file also housed MemoryTrackerTest, whose four unit tests
(trackAndReport, trackCopies, clearResetsState, fileBackedTracking)
cover the live MemoryTracker aggregate reports. Restored verbatim in
its own file; only the planner-owned imports are dropped.

Part of #1146's memory-accounting groundwork.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The creation path learned to consult ExecutionContext.memoryScope in
#1145; op outputs still reached the GC because DefaultCpuOps allocates
them through its TensorDataFactory. This closes that half (#1146):

- TensorDataFactory.adoptFloatArray — the op-output entry point. Unlike
  wrapFloatArray, the buffer is a fresh, ownership-transferred result,
  which licenses a factory to relocate it. DenseTensorDataFactory adopts
  FP32/FP16 zero-copy (one copy FEWER than before on the common path,
  which paid fromFloatArray's copyOf); narrow dtypes keep their tagged
  fromFloatArray path.
- ScopedTensorDataFactory decorates any factory: under a non-Ambient
  scope, dense-FP32 zeros/ones/full/init/fromFloatArray/adoptFloatArray
  land in the slab as StorageFloatTensorData; the region is always fully
  written because a reset slab is dirty. wrap* stays un-intercepted —
  its zero-copy caller-owned contract must not die at reset().
- ExecutionContext.withTensorDataFactory + the DirectCpuExecutionContext
  override let ScopedExecutionContext rebuild the base around the scoped
  factory and re-bind created tensors to the rebuilt ops, so 'a + b'
  dispatches into scope-allocating ops.
- The FP32 fast paths are offset-aware (floatWindowOf): slab-backed
  operands stay on the flat primitive loops instead of falling to the
  boxed generic path (#949's 83%-overhead cliff — live since #1145 for
  scope-created tensors, now closed). Tail sites read slab data as its
  exact logical window through floatBufferOf (one copy, still far off
  the boxed path). DefaultCpuOpsJvm's 18 direct constructions route
  through floatResult; its Panama vector paths fall back to the common
  window loops for slab operands (follow-up: offset-aware vectors).
- KernelDispatch.matmul receives dispatchScope(), so requantize/prepack
  adapter allocations land in the slab too.
- ScopedOpOutputsTest pins: slab-backed outputs, flat peakFloats across
  8 steps with zero overflow (steady-state decode allocates no new slab
  bytes), bit-identical numerics vs Ambient, loud StorageClosedException
  on stale reads, and the overflow path.

Also restores MemoryTrackerTest (lost with MemoryPlannerTest.kt in #1142).

Closes #1146.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1168 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 739a260 into develop Aug 26, 2026
17 of 18 checks passed
@michalharakal
michalharakal deleted the feature/1146-op-outputs-through-scope branch August 26, 2026 12:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Route eager op outputs through ctx.memoryScope (DefaultCpuOps + decode-loop adoption)

1 participant