Footprint capability: header-only planInput for safetensors/ONNX, external_data fix, EDGE profile - #1170
Merged
Conversation
…ly planInput, external_data fix, EDGE profile There was no quick way to check whether a model fits on an embedded device with limited memory and compute before investing a day in converting it. This lands the library capability (#1169); the multi-format CLI over it incubates in SKaiNET-research until a stable core release carries these APIs. - StreamingSafeTensorsReader.planInput / StreamingOnnxReader.planInput: PlanInput with null geometry — weights-only plans through the same render/verdict/suggestion pipeline GGUF uses. Byte counts are authoritative: safetensors from data_offsets, ONNX from raw_data. - ONNX external_data (field 13) is parsed instead of skipped: >2 GB models keep their weights in a sibling file, and exactly those models previously reported ~0 bytes — a fit verdict that lied where it mattered most. Sizes are Long throughout (estimatedBytesLong; the Int view clamps instead of wrapping negative). - PlannerProfile.EDGE: an embedded device where the number the caller passes IS the usable RAM — reserve deliberately zero and documented to stay zero; weights mapped, KV auto-quantized past 80%. - SafeTensorsDataTypeMapper no longer printlns a WARNING into stdout mid-parse; UNKNOWN is the answer. - Tests pin the capability so later memory-layout refactors cannot silently lose it: per-format planInput sums, the 3 GiB external tensor (Long correctness), unknown-dtype pricing via offsets. Deferred (noted in #1169): HF config.json geometry, sharded index support, /proc/meminfo DeviceMemory provider, ONNX TensorId maps. Part of #1169. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
michalharakal
force-pushed
the
feature/1169-model-footprint
branch
from
August 26, 2026 13:25
939e2e3 to
e6280d1
Compare
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #1169 — the library capability; the multi-format CLI over it incubates in SKaiNET-research (
feature/model-footprint-cli) until a stable core release carries these APIs. Companion write-up: SKaiNET-developers/SKaiNET-research#5.There was no quick way to check whether a model fits on an embedded device with limited memory and compute before investing a day in converting it. This PR makes the answer computable from headers/metadata alone, for all three formats:
StreamingSafeTensorsReader.planInput/StreamingOnnxReader.planInput—PlanInputwithgeometry = null: weights-only plans through the same render/verdict/suggestion pipeline GGUF uses. Byte counts are authoritative (safetensorsdata_offsets, ONNXraw_data).external_dataparsed instead of skipped — >2 GB models keep weights in a sibling file, and exactly those models previously reported ~0 bytes: a fit verdict that lied where it mattered most. Sizes areLongthroughout (estimatedBytesLong; theIntview clamps instead of wrapping negative). Pinned by a 3 GiB-tensor test.PlannerProfile.EDGE— an embedded device where the number the caller passes IS the usable RAM: reserve deliberately zero and documented to stay zero; weights mapped, KV auto-quantized past 80 %.SafeTensorsDataTypeMapperno longerprintlns a WARNING into stdout mid-parse.planInput→EDGE.plan(...)→ verdict).Why the tests matter beyond this PR: they pin the memory-counting capability itself — per-format sums, the 3 GiB
Long-correctness case, unknown-dtype pricing via offsets — so later memory-layout refactors cannot silently lose it.Full pr-gate green (before the CLI split; the split removes app-level code only, and the affected module suites re-ran green).
🤖 Generated with Claude Code