Skip to content

P7: AllocationResolver + plan→load wiring (re-land #1153/#1154 onto develop) - #1155

Merged
michalharakal merged 2 commits into
developfrom
feature/1144-plan-load-wiring
Aug 26, 2026
Merged

michalharakal merged 2 commits into
developfrom
feature/1144-plan-load-wiring

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Re-lands #1153 + #1154 against develop.

Both PRs were merged, but into their stacked base branches (feature/1142-…, feature/1143-…) rather than develop — GitHub only retargets a stacked PR automatically when its base branch is deleted at merge, and the bases weren't deleted. So the AllocationResolver (#1143) and plan→load wiring (#1144) never reached develop. This PR is those same two commits, rebased onto current develop (post-#1151/#1152); see the original PRs for the full descriptions and the green full-gate runs.

Tip for the rest of the stack: ticking "delete branch" when merging makes GitHub retarget the child PR to develop automatically.

🤖 Generated with Claude Code

michalharakal and others added 2 commits August 26, 2026 10:56
Placement joins WeightForm as a resolved decision (#1133's answer):
AllocationResolver is a pure function of what will be held (the resolved
form), what the profile says (domainFor and its threshold) and what the
platform can do (StorageCapabilities, injectable so a test can resolve
for a platform it is not running on).

PlanTensor.allocation stops hardcoding MMAP_FILE/MODEL — a spec that
claimed every weight was mapped even on a platform that cannot map, and
even for a dequantized copy that no longer matches the file. Mapping now
requires all three: the form asks MAPPED, the platform can map, and the
bytes really are the file's bytes; everything else falls to the
profile's heap/off-heap threshold over the bytes actually held.

AllocationResolver.explain() renders the decision with its reason, so a
plan can say where every tensor landed and why.

Closes #1143.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…iring

Plan and load were two disjoint pipelines sharing vocabulary and never
talking: nothing on the load path consulted what the resolvers decided.

StreamingGgufParametersLoader gains weightFormFor — per-tensor forms
with an explicit, documented precedence: your function > the uniform
weightForm > the three legacy parameters. The user-wins channel is
named as such: whatever you pass outranks every resolver, including
'everything dense on the managed heap'.

ResolvedGguf ties it together: header-only planInput, forms resolved
from file × profile × kernel capability, overrides applied, and a
loader that delivers exactly those forms. The same resolved input is
priceable (profiledPlan().requireFits() refuses pre-load) and
explainable (explainPlacements() — one line per weight, where and why)
before a byte of payload is read.

Closes #1144.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit f3c88f1 into develop Aug 26, 2026
14 of 15 checks passed
@michalharakal
michalharakal deleted the feature/1144-plan-load-wiring branch August 26, 2026 08:58
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1155 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant