feat(io): report every weight the loader re-encodes on the way in - #1123
Merged
Merged
Conversation
Closes #1117. Slice 4 of #1109. "Why is this model bigger than the file" had no answer except whoever remembered which policy was set. Now each conversion the loader performs emits a TraceEvent.AdapterInserted naming the tensor and the sizes either side; a weight that arrives in the form it will be used in emits nothing. Reusing AdapterInserted rather than adding an event type, because its own documentation already describes this — "the dispatcher inserted a conversion (dequantize, requantize, gather)" — and a loader-side re-encode is the same thing at a different moment. It also means the Perfetto, JFR and android.os.Trace exporters need no changes to show it, which is the third acceptance criterion satisfied by not writing code. The event gained one defaulted field, bytesBefore, and a bytesDelta derived from it: a conversion's cost is the difference between two sizes and one number cannot express it. Perfetto and JFR now carry both. Three conversions emit, not one. The issue asked for the ones the resolver causes, which today means form-driven dequantization alone — but the loader performs two widenings nobody ever asked for, and those are the ones most worth seeing. Ternary tensors widen to FP32 whatever the policy, because packed ternary storage does not exist yet (#1033): roughly 20×, and no flag on this loader turns it off. Narrow floats double unless keepF16Native or keepBf16Native is set. Reporting only the requested conversion would answer the easy half of the question. The sink defaults to NoopTraceSink and every emission is guarded on isEnabled, so a caller who asked for no trace builds no Formats and allocates nothing. Gate: scripts/pr-gate.sh — all legs passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
aharakal
approved these changes
Aug 25, 2026
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1117. Slice 4 of 5 for #1109.
"Why is this model bigger than the file it came from" had no answer except whoever remembered which policy was set. Now each conversion emits a
TraceEvent.AdapterInsertednaming the tensor and the sizes either side — and a weight that arrives in the form it will be used in emits nothing at all.Reusing the event rather than adding one
AdapterInserted's own documentation already describes this: "the dispatcher inserted a conversion (dequantize, requantize, gather)". A loader-side re-encode is the same thing at a different moment. Reusing it means the Perfetto, JFR andandroid.os.Traceexporters need no changes to show it — the third acceptance criterion satisfied by not writing code.It gained one defaulted field,
bytesBefore, plus a derivedbytesDelta. A conversion's cost is the difference between two sizes, and one number cannot express it; Perfetto and JFR now carry both.Three conversions, not one
The issue asked for the conversions the resolver causes, which today means form-driven dequantization alone. But the loader performs two other widenings that nobody ever asked for, and those are the ones most worth seeing:
dequantize-on-loadDequantizeTo(FP32)widen-ternary-no-packed-storagewiden-f16/widen-bf16keepF16Native/keepBf16NativeReporting only the conversion someone explicitly requested would answer the easy half of the question. The ternary one in particular is invisible today: it is not a policy, it is a gap, and the trace is where a gap that costs 20× should show up.
Cost when nobody is listening
The sink defaults to
NoopTraceSink, and every emission is guarded onisEnabled, so a caller who asked for no trace builds noFormats and allocates nothing.The tensor is named by parsing the GGUF name as a
TensorId—blk.0.attn_q.weightis already dotted and renders back as itself, so no new field was needed to carry a name.Tests
Five: the silent case is genuinely silent; a dequantized weight reports the right tensor, both sizes and the delta; the ternary widening is caught with its multiple asserted; a narrow float reports its doubling; and the default loader records nothing.
Gate
scripts/pr-gate.sh— all legs passed.🤖 Generated with Claude Code