fix(io-gguf): stream GGUF dequantization — no boxed payloads, no defensive copies (#782) - #965
Merged
Merged
Conversation
…nsive copies (#782) Loading a 1.1B Q4_K_M via DEQUANTIZE_TO_FP32 transiently needed >12 GB heap against a ~4.4 GB dense-FP32 floor. Three compounding causes, all in skainet-io-gguf: 1. GGUFReader (legacy in-memory reader) eagerly materialized EVERY tensor payload as a boxed List<Any> at parse time via readDataByType — measured at 41x the payload size in allocations (copyOfRange + one boxed element per byte + chunked() garbage): ~26 GB for a 637 MB file. Payloads are now constant-space lazy views that decode elements on access. Same List<Any> static type, same contents, same equality — only the storage strategy changed. 2. StreamingGgufParametersLoader paid the tensor factory's full-size defensive copyOf for every dense tensor on top of the dequant intermediate. The loader owns every array it decodes, so it now wraps them zero-copy (ctx.wrapFloatArray). 3. The K-quant kernels (Q4_K/Q5_K/Q6_K) allocated per-block copyOfRange scratch; they now index the source buffer directly, so a full-tensor dequant allocates exactly the destination FloatArray. StreamingGgufParametersLoader also gains an optional quantPolicy parameter (default NATIVE_OPTIMIZED = the historical packed behavior, bit-for-bit). DEQUANTIZE_TO_FP32 streams each quantized tensor block-by-block straight into its destination array and wraps it — peak transient per tensor is the packed source bytes. RAW_BYTES is rejected eagerly. Measured (heap-instrumented tests, per-thread allocation counters): - historical copy chain: 2.1-2.3x the FP32 total in allocations - fixed streaming DEQUANTIZE_TO_FP32: 1.38x allocated, peak live ~1.05x - legacy reader parse: 2.0x file size (was 41x the payload) New tests: StreamingDequantPolicyParityTest pins the dequant path to the packed block accessors bit-exactly across all seven supported formats (multi-block, pseudo-random payloads); DequantHeapUsageTest enforces the transient-allocation budget and documents the historical numbers. Closes #782 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Aug 11, 2026
aharakal
previously approved these changes
Aug 11, 2026
aharakal
approved these changes
Aug 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the GGUF
DEQUANTIZE_TO_FP32transient over-allocation: a 1.1B Q4_K_M needed >12 GB heap against a ~4.4 GB dense-FP32 floor (and the 270M follow-up measurement showed the same ~3x ratio on the streaming path). All three compounding causes live inskainet-io-gguf:Root causes
GGUFReader) — withloadTensorData = true(the default), every tensor payload was eagerly materialized as a boxedList<Any>at parse time throughreadDataByType(copyOfRange+ one boxed element per byte +chunked()garbage). Measured with per-thread allocation counters: 41x the payload size — ~26 GB of allocation for a 637 MB file. This is the ">12 GB / needs 32 GB" path, and the "boxed Float" the issue body pinpoints.StreamingGgufParametersLoader) — every dense tensor paid the tensor factory's full-sizecopyOf(DenseTensorDataFactory.createFloatTensorData) on top of the dequant intermediate: 2x the FP32 size in transients per tensor.dequantQ4K/Q5K/Q6KFromBytesallocatedcopyOfRangescratch for every 144–210-byte block (~1.2x the packed size in GC churn per tensor).Fix (behavior-identical)
GGUFReaderpayloads become constant-space lazy views — sameList<Any>static type, same contents, sameAbstractListequality; elements decode on access. Parse-time allocation drops from 41x payload to ~2x file size.StreamingGgufParametersLoaderwraps its loader-owned arrays zero-copy (ctx.wrapFloatArray) instead of paying the defensive copy — it owns every array it decodes.FloatArray.quantPolicyparameter onStreamingGgufParametersLoader(defaultNATIVE_OPTIMIZED= the historical packed behavior, bit-for-bit).DEQUANTIZE_TO_FP32streams each quantized tensor block-by-block straight into its destination array and wraps it — peak transient per tensor = the packed source bytes, exactly the issue's suggested fix.RAW_BYTESis rejected eagerly with guidance.Before / after (heap-instrumented tests, per-thread allocation counters)
DEQUANTIZE_TO_FP32(after)For the issue's TinyLlama numbers that means the eager FP32 materialization fits in ≈ the 4.4 GB dense floor + the 637 MB packed source + per-tensor slack — the "~5–6 GB, not >12 GB" target.
Tests
StreamingDequantPolicyParityTest— pins the dequant path to the packed block accessors bit-exactly across all seven supported quant formats (multi-block, pseudo-random payloads; single-block tensors can pass by accident), asserts the default policy still deliversPackedBlockStorage, and the eagerRAW_BYTESrejection.DequantHeapUsageTest— synthetic multi-tensor GGUF (Q4_K/Q6_K/Q8_0/F16/F32, 42 MB dense): asserts the fixed path's transient allocation stays within budget (measured 1.38x total, well under the issue's 1.2x-of-floor acceptance bar for peak live), documents the historical 2.1–2.3x chain, and gates the legacy reader's parse allocation at O(file size) with a lazy-view content spot-check.:skainet-io:skainet-io-ggufsuite: jvmTest 109/109, linuxX64Test 21/21 green. (jsBrowserTestneeds a local browser install; not run here.)apiCheck:skainet-io-ggufhas no BCV dump; the tracked modules were run —skainet-lang-core:jvmApiCheckfails identically on pristine develop (pre-existing drift from the earlierDim/dynamic-axes additions, byte-identical diff with and without this PR; same forskainet-compile-hlo). This PR adds zero API drift.Notes
SKaiNET-transformersloaders (DecoderGgufWeightLoader.streamingTensorToTensorand thefromGguf(Source)legacy overload) still run their own dequant-then-copy sequence; they inherit causes 1 and 3 fixed here transparently and can drop their remaining intermediate by adoptingquantPolicy = DEQUANTIZE_TO_FP32/wrapFloatArrayon the next engine pin bump — tracked as part of the W1 rollout.Closes #782
🤖 Generated with Claude Code