Skip to content

feat(io): report every weight the loader re-encodes on the way in - #1123

Merged
michalharakal merged 1 commit into
developfrom
feature/1117-weight-form-trace
Aug 25, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1117-weight-form-trace

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1117. Slice 4 of 5 for #1109.

"Why is this model bigger than the file it came from" had no answer except whoever remembered which policy was set. Now each conversion emits a TraceEvent.AdapterInserted naming the tensor and the sizes either side — and a weight that arrives in the form it will be used in emits nothing at all.

Reusing the event rather than adding one

AdapterInserted's own documentation already describes this: "the dispatcher inserted a conversion (dequantize, requantize, gather)". A loader-side re-encode is the same thing at a different moment. Reusing it means the Perfetto, JFR and android.os.Trace exporters need no changes to show it — the third acceptance criterion satisfied by not writing code.

It gained one defaulted field, bytesBefore, plus a derived bytesDelta. A conversion's cost is the difference between two sizes, and one number cannot express it; Perfetto and JFR now carry both.

Three conversions, not one

The issue asked for the conversions the resolver causes, which today means form-driven dequantization alone. But the loader performs two other widenings that nobody ever asked for, and those are the ones most worth seeing:

kind when cost
dequantize-on-load the resolved form says DequantizeTo(FP32) Q4_K → FP32, ~8×
widen-ternary-no-packed-storage always, for TQ1_0/TQ2_0 — packed ternary storage does not exist yet (#1033) ~20×, and no flag on this loader turns it off
widen-f16 / widen-bf16 unless keepF16Native / keepBf16Native 2×

Reporting only the conversion someone explicitly requested would answer the easy half of the question. The ternary one in particular is invisible today: it is not a policy, it is a gap, and the trace is where a gap that costs 20× should show up.

Cost when nobody is listening

The sink defaults to NoopTraceSink, and every emission is guarded on isEnabled, so a caller who asked for no trace builds no Formats and allocates nothing.

The tensor is named by parsing the GGUF name as a TensorId — blk.0.attn_q.weight is already dotted and renders back as itself, so no new field was needed to carry a name.

Tests

Five: the silent case is genuinely silent; a dequantized weight reports the right tensor, both sizes and the delta; the ternary widening is caught with its multiple asserted; a narrow float reports its doubling; and the default loader records nothing.

Gate

scripts/pr-gate.sh — all legs passed.

🤖 Generated with Claude Code

Closes #1117. Slice 4 of #1109.

"Why is this model bigger than the file" had no answer except whoever
remembered which policy was set. Now each conversion the loader performs
emits a TraceEvent.AdapterInserted naming the tensor and the sizes either
side; a weight that arrives in the form it will be used in emits nothing.

Reusing AdapterInserted rather than adding an event type, because its own
documentation already describes this — "the dispatcher inserted a conversion
(dequantize, requantize, gather)" — and a loader-side re-encode is the same
thing at a different moment. It also means the Perfetto, JFR and
android.os.Trace exporters need no changes to show it, which is the third
acceptance criterion satisfied by not writing code.

The event gained one defaulted field, bytesBefore, and a bytesDelta derived
from it: a conversion's cost is the difference between two sizes and one
number cannot express it. Perfetto and JFR now carry both.

Three conversions emit, not one. The issue asked for the ones the resolver
causes, which today means form-driven dequantization alone — but the loader
performs two widenings nobody ever asked for, and those are the ones most
worth seeing. Ternary tensors widen to FP32 whatever the policy, because
packed ternary storage does not exist yet (#1033): roughly 20×, and no flag
on this loader turns it off. Narrow floats double unless keepF16Native or
keepBf16Native is set. Reporting only the requested conversion would answer
the easy half of the question.

The sink defaults to NoopTraceSink and every emission is guarded on
isEnabled, so a caller who asked for no trace builds no Formats and
allocates nothing.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal requested a review from aharakal August 25, 2026 14:47
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1123 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 4add306 into develop Aug 25, 2026
16 checks passed
@michalharakal
michalharakal deleted the feature/1117-weight-form-trace branch August 25, 2026 14:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Every repack and dequant says so on the trace

2 participants