fix(io): map all seven GGUF quant formats to their TensorEncodings; real byte counts for Opaque (#928) - #936
Merged
Merged
Conversation
…eal byte counts for Opaque StreamingGGUFReader.ggmlTypeToEncoding mapped only Q4_K and Q8_0 to dedicated encodings. Q4_0, Q5_0, Q5_1, Q5_K and Q6_K — all of which have TensorEncoding objects with correct block sizes and CPU kernels — fell through to Opaque(name, 0), with a comment promising 'size computed from tensor info' that nothing ever computed. Since Opaque.physicalBytes returns rawBytes non-null, TensorStorage.physicalBytes never fell back to buffer.sizeInBytes: every such tensor reported 0 physical bytes, memory reports were wrong for five of seven supported formats, and compressionRatio silently returned 1.0. Map all seven quant formats to their encodings, and pass the tensor's real nBytes into the Opaque fallback for genuinely unknown types, so physicalBytes stays truthful either way. Both loadTensorStorage and loadTensorStorageMapped call sites updated. Tests: synthesized GGUF with a Q4_0 tensor (dedicated encoding, 18-byte block) and a Q8_1 tensor (no dedicated encoding -> Opaque with real 40-byte count), covering heap and mapped storage paths. Closes #928
aharakal
previously approved these changes
Aug 10, 2026
aharakal
previously approved these changes
Aug 10, 2026
aharakal
approved these changes
Aug 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Full analysis in #928.
ggmlTypeToEncodingmapped only Q4_K/Q8_0; Q4_0, Q5_0, Q5_1, Q5_K and Q6_K — which all haveTensorEncodingobjects and CPU kernels — fell toOpaque(name, 0)(with a comment promising "size computed from tensor info" that nothing computed). BecauseOpaque.physicalBytesreturns its stored count non-null,TensorStorage.physicalBytesnever falls back tobuffer.sizeInBytes: five of seven supported quant formats reported 0 physical bytes, andcompressionRatiosilently returned 1.0.Fix: all seven formats map to their dedicated encodings; genuinely unknown types now carry the tensor's real
nBytesinOpaque, sophysicalBytesstays truthful either way. BothloadTensorStorageandloadTensorStorageMappedpass the size.Side benefit: encodings travel to StableHLO as the
skainet.tensor_encodingsmodule attribute, so the compile path now sees real layouts for these formats too.Tests: new
StreamingGGUFReaderEncodingTest— synthesized GGUF with a Q4_0 tensor (dedicated encoding, 18-byte block asserted) and a Q8_1 tensor (Opaque with the real 40-byte count asserted,physicalBytes != 0), heap and mapped paths.Verified:
:skainet-io:skainet-io-gguf:jvmTestgreen (full module).Note: CHANGELOG entry included — trivial
[Unreleased]conflicts expected with the sibling #927–#931 PRs.Closes #928