Skip to content

Eager lifetime: wire ExecutionContext.memoryScope into tensor creation - #1156

Merged
michalharakal merged 1 commit into
developfrom
feature/1145-memoryscope-wiring
Aug 26, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1145-memoryscope-wiring

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1145 (parent #1135). Stacked on #1155 — merge that first (and tick "delete branch" so this one retargets to develop automatically).

The answer to #1135, wired: eager needs no graph to reuse buffers — the Scope split already models it, and ForwardScope.reset() at step boundaries replaces graph liveness analysis. What was missing was any reader of ExecutionContext.memoryScope. This is the reader.

  • zeros/ones/full/fromFloatArray consult memoryScope when it is not Scope.Ambient and the dtype is dense FP32: bytes come from the scope's slab as the new StorageFloatTensorData, die at reset(), and a use-after-reset is a loud StorageClosedException. Ambient — the default everywhere — short-circuits to the factory path untouched.
  • ScopedExecutionContext + ctx.forwardScope(slabFloats) { scoped, scope -> … } is the opt-in: steady-state decode allocates zero new slab bytes per step (pinned by test).
  • StorageFloatTensorData deliberately does not implement FloatArrayTensorData: slab slices have nonzero arrayOffset and the ops fast paths assume offset 0 — the top silent-breakage risk. A backend-cpu parity test runs real CPU matmul/add on slab-backed tensors at nonzero offsets, across resets, against Ambient results.
  • ScratchPool stays, unmerged: intra-kernel workspace vs inter-op activation lifetime are different layers, now documented on both properties.
  • Op outputs still allocate raw arrays — routing DefaultCpuOps' 34 sites through the scope plus decode-loop adoption is Route eager op outputs through ctx.memoryScope (DefaultCpuOps + decode-loop adoption) #1146 (filed).

Full pr-gate green on the rebased stack.

🤖 Generated with Claude Code

The answer to #1135: eager needs no graph to reuse buffers — the Scope
split already models it. Eager lifetimes follow call structure, so
ForwardScope.reset() at step boundaries replaces graph liveness
analysis. What was missing was any reader of memoryScope; this is the
reader.

zeros/ones/full/fromFloatArray consult memoryScope when it is not
Scope.Ambient and the dtype is dense FP32: the bytes come from the
scope's slab as StorageFloatTensorData, die at reset(), and a
use-after-reset is a StorageClosedException naming the storage.
Ambient — the default everywhere — short-circuits to the factory path
untouched.

StorageFloatTensorData deliberately does not implement
FloatArrayTensorData: a slab slice has a nonzero arrayOffset and the
ops fast paths that unwrap 'buffer' assume offset 0. The backend-cpu
parity test pins real CPU ops on slab-backed tensors at nonzero
offsets against the Ambient results, across resets.

ScratchPool stays, unmerged: intra-kernel workspace vs inter-op
activation lifetime are different layers, now documented on both.

Op *outputs* still allocate raw arrays; routing them through the scope
is #1146.

Closes #1145.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Base automatically changed from feature/1144-plan-load-wiring to develop August 26, 2026 08:58
@michalharakal
michalharakal merged commit eb96541 into develop Aug 26, 2026
5 checks passed
@michalharakal
michalharakal deleted the feature/1145-memoryscope-wiring branch August 26, 2026 09:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eager lifetime: wire ExecutionContext.memoryScope into tensor creation

1 participant