Skip to content

fix(io): map all seven GGUF quant formats to their TensorEncodings; real byte counts for Opaque (#928) - #936

Merged
michalharakal merged 6 commits into
developfrom
fix/gguf-encoding-map-928
Aug 10, 2026
Merged

michalharakal merged 6 commits into
developfrom
fix/gguf-encoding-map-928

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Full analysis in #928.

ggmlTypeToEncoding mapped only Q4_K/Q8_0; Q4_0, Q5_0, Q5_1, Q5_K and Q6_K — which all have TensorEncoding objects and CPU kernels — fell to Opaque(name, 0) (with a comment promising "size computed from tensor info" that nothing computed). Because Opaque.physicalBytes returns its stored count non-null, TensorStorage.physicalBytes never falls back to buffer.sizeInBytes: five of seven supported quant formats reported 0 physical bytes, and compressionRatio silently returned 1.0.

Fix: all seven formats map to their dedicated encodings; genuinely unknown types now carry the tensor's real nBytes in Opaque, so physicalBytes stays truthful either way. Both loadTensorStorage and loadTensorStorageMapped pass the size.

Side benefit: encodings travel to StableHLO as the skainet.tensor_encodings module attribute, so the compile path now sees real layouts for these formats too.

Tests: new StreamingGGUFReaderEncodingTest — synthesized GGUF with a Q4_0 tensor (dedicated encoding, 18-byte block asserted) and a Q8_1 tensor (Opaque with the real 40-byte count asserted, physicalBytes != 0), heap and mapped paths.

Verified: :skainet-io:skainet-io-gguf:jvmTest green (full module).

Note: CHANGELOG entry included — trivial [Unreleased] conflicts expected with the sibling #927–#931 PRs.

Closes #928

…eal byte counts for Opaque

StreamingGGUFReader.ggmlTypeToEncoding mapped only Q4_K and Q8_0 to
dedicated encodings. Q4_0, Q5_0, Q5_1, Q5_K and Q6_K — all of which have
TensorEncoding objects with correct block sizes and CPU kernels — fell
through to Opaque(name, 0), with a comment promising 'size computed from
tensor info' that nothing ever computed. Since Opaque.physicalBytes
returns rawBytes non-null, TensorStorage.physicalBytes never fell back to
buffer.sizeInBytes: every such tensor reported 0 physical bytes, memory
reports were wrong for five of seven supported formats, and
compressionRatio silently returned 1.0.

Map all seven quant formats to their encodings, and pass the tensor's
real nBytes into the Opaque fallback for genuinely unknown types, so
physicalBytes stays truthful either way. Both loadTensorStorage and
loadTensorStorageMapped call sites updated.

Tests: synthesized GGUF with a Q4_0 tensor (dedicated encoding, 18-byte
block) and a Q8_1 tensor (no dedicated encoding -> Opaque with real
40-byte count), covering heap and mapped storage paths.

Closes #928
@michalharakal
michalharakal requested a review from aharakal August 10, 2026 08:54
aharakal
aharakal previously approved these changes Aug 10, 2026
aharakal
aharakal previously approved these changes Aug 10, 2026
@michalharakal
michalharakal merged commit 8f9392b into develop Aug 10, 2026
9 of 10 checks passed
@michalharakal
michalharakal deleted the fix/gguf-encoding-map-928 branch August 10, 2026 13:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

StreamingGGUFReader maps only Q4_K/Q8_0 to real TensorEncodings — Q6_K/Q5_K/Q4_0/Q5_0/Q5_1 become Opaque(…, 0) and corrupt every memory report

2 participants