Skip to content

fix(io-gguf): stream GGUF dequantization — no boxed payloads, no defensive copies (#782) - #965

Merged
michalharakal merged 2 commits into
developfrom
fix/782-dequant-overalloc
Aug 11, 2026
Merged

michalharakal merged 2 commits into
developfrom
fix/782-dequant-overalloc

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Fixes the GGUF DEQUANTIZE_TO_FP32 transient over-allocation: a 1.1B Q4_K_M needed >12 GB heap against a ~4.4 GB dense-FP32 floor (and the 270M follow-up measurement showed the same ~3x ratio on the streaming path). All three compounding causes live in skainet-io-gguf:

Root causes

  1. Whole-file boxed materialization (legacy GGUFReader) — with loadTensorData = true (the default), every tensor payload was eagerly materialized as a boxed List<Any> at parse time through readDataByType (copyOfRange + one boxed element per byte + chunked() garbage). Measured with per-thread allocation counters: 41x the payload size — ~26 GB of allocation for a 637 MB file. This is the ">12 GB / needs 32 GB" path, and the "boxed Float" the issue body pinpoints.
  2. Defensive copy on the load path (StreamingGgufParametersLoader) — every dense tensor paid the tensor factory's full-size copyOf (DenseTensorDataFactory.createFloatTensorData) on top of the dequant intermediate: 2x the FP32 size in transients per tensor.
  3. Per-block scratch in the K-quant kernels — dequantQ4K/Q5K/Q6KFromBytes allocated copyOfRange scratch for every 144–210-byte block (~1.2x the packed size in GC churn per tensor).

Fix (behavior-identical)

  • GGUFReader payloads become constant-space lazy views — same List<Any> static type, same contents, same AbstractList equality; elements decode on access. Parse-time allocation drops from 41x payload to ~2x file size.
  • StreamingGgufParametersLoader wraps its loader-owned arrays zero-copy (ctx.wrapFloatArray) instead of paying the defensive copy — it owns every array it decodes.
  • The K-quant kernels index the source buffer directly — a full-tensor dequant now allocates exactly the destination FloatArray.
  • New optional quantPolicy parameter on StreamingGgufParametersLoader (default NATIVE_OPTIMIZED = the historical packed behavior, bit-for-bit). DEQUANTIZE_TO_FP32 streams each quantized tensor block-by-block straight into its destination array and wraps it — peak transient per tensor = the packed source bytes, exactly the issue's suggested fix. RAW_BYTES is rejected eagerly with guidance.

Before / after (heap-instrumented tests, per-thread allocation counters)

Path Allocation vs dense-FP32 total
legacy reader payload materialization (before) 41x the payload (~26 GB for TinyLlama's 637 MB)
historical streaming chain: dequant intermediate + factory copy (before) 2.1–2.3x
fixed streaming DEQUANTIZE_TO_FP32 (after) 1.38x allocated, peak live ≈ 1.05x
legacy reader parse (after) 2.0x file size (the file buffer + metadata)

For the issue's TinyLlama numbers that means the eager FP32 materialization fits in ≈ the 4.4 GB dense floor + the 637 MB packed source + per-tensor slack — the "~5–6 GB, not >12 GB" target.

Tests

  • StreamingDequantPolicyParityTest — pins the dequant path to the packed block accessors bit-exactly across all seven supported quant formats (multi-block, pseudo-random payloads; single-block tensors can pass by accident), asserts the default policy still delivers PackedBlockStorage, and the eager RAW_BYTES rejection.
  • DequantHeapUsageTest — synthetic multi-tensor GGUF (Q4_K/Q6_K/Q8_0/F16/F32, 42 MB dense): asserts the fixed path's transient allocation stays within budget (measured 1.38x total, well under the issue's 1.2x-of-floor acceptance bar for peak live), documents the historical 2.1–2.3x chain, and gates the legacy reader's parse allocation at O(file size) with a lazy-view content spot-check.
  • Full :skainet-io:skainet-io-gguf suite: jvmTest 109/109, linuxX64Test 21/21 green. (jsBrowserTest needs a local browser install; not run here.)
  • apiCheck: skainet-io-gguf has no BCV dump; the tracked modules were run — skainet-lang-core:jvmApiCheck fails identically on pristine develop (pre-existing drift from the earlier Dim/dynamic-axes additions, byte-identical diff with and without this PR; same for skainet-compile-hlo). This PR adds zero API drift.

Notes

  • Downstream SKaiNET-transformers loaders (DecoderGgufWeightLoader.streamingTensorToTensor and the fromGguf(Source) legacy overload) still run their own dequant-then-copy sequence; they inherit causes 1 and 3 fixed here transparently and can drop their remaining intermediate by adopting quantPolicy = DEQUANTIZE_TO_FP32 / wrapFloatArray on the next engine pin bump — tracked as part of the W1 rollout.
  • SKEEP-003 alignment: this is the design doc's first implementation slice (staging stage of the IO pipeline improvement) — see docs(skeep): SKEEP-003 amendment — deprecate-don't-delete migration rule + implementation slices (#782, #921) #963. The fix is deliberately conservative: additive parameter with historical default, implementation swaps behind unchanged types, no public API removed.

Closes #782

🤖 Generated with Claude Code

…nsive copies (#782)

Loading a 1.1B Q4_K_M via DEQUANTIZE_TO_FP32 transiently needed >12 GB
heap against a ~4.4 GB dense-FP32 floor. Three compounding causes, all in
skainet-io-gguf:

1. GGUFReader (legacy in-memory reader) eagerly materialized EVERY tensor
   payload as a boxed List<Any> at parse time via readDataByType —
   measured at 41x the payload size in allocations (copyOfRange + one
   boxed element per byte + chunked() garbage): ~26 GB for a 637 MB file.
   Payloads are now constant-space lazy views that decode elements on
   access. Same List<Any> static type, same contents, same equality —
   only the storage strategy changed.

2. StreamingGgufParametersLoader paid the tensor factory's full-size
   defensive copyOf for every dense tensor on top of the dequant
   intermediate. The loader owns every array it decodes, so it now wraps
   them zero-copy (ctx.wrapFloatArray).

3. The K-quant kernels (Q4_K/Q5_K/Q6_K) allocated per-block copyOfRange
   scratch; they now index the source buffer directly, so a full-tensor
   dequant allocates exactly the destination FloatArray.

StreamingGgufParametersLoader also gains an optional quantPolicy
parameter (default NATIVE_OPTIMIZED = the historical packed behavior,
bit-for-bit). DEQUANTIZE_TO_FP32 streams each quantized tensor
block-by-block straight into its destination array and wraps it — peak
transient per tensor is the packed source bytes. RAW_BYTES is rejected
eagerly.

Measured (heap-instrumented tests, per-thread allocation counters):
- historical copy chain: 2.1-2.3x the FP32 total in allocations
- fixed streaming DEQUANTIZE_TO_FP32: 1.38x allocated, peak live ~1.05x
- legacy reader parse: 2.0x file size (was 41x the payload)

New tests: StreamingDequantPolicyParityTest pins the dequant path to the
packed block accessors bit-exactly across all seven supported formats
(multi-block, pseudo-random payloads); DequantHeapUsageTest enforces the
transient-allocation budget and documents the historical numbers.

Closes #782

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
aharakal
aharakal previously approved these changes Aug 11, 2026
@michalharakal
michalharakal merged commit c1d3248 into develop Aug 11, 2026
11 checks passed
@michalharakal
michalharakal deleted the fix/782-dequant-overalloc branch August 11, 2026 15:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GGUF DEQUANTIZE_TO_FP32 over-allocates: 1.1B Q4_K_M needs >12 GB heap transiently (~4.4 GB legit)

2 participants