Skip to content

fix(io): fail fast on unsupported GGUF tensor types; load Q4_0/Q5_0/Q5_1 packed (#919) - #925

Merged
michalharakal merged 2 commits into
developfrom
fix/gguf-loader-fail-fast-919
Aug 10, 2026
Merged

michalharakal merged 2 commits into
developfrom
fix/gguf-loader-fail-fast-919

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

What

StreamingGgufParametersLoader silently skipped any tensor type outside its when — a SKIP progress string was the only signal, the load "succeeded" with missing weights, and the failure surfaced far away in graph build / the first forward pass (the load-time half of #654). The deprecated legacy GgufParametersLoader already hard-failed here; the streaming loader regressed that. Full analysis in #919.

Two distinct gaps hid in that else:

  1. Truly unsupported types (e.g. Q4_1) → now an eager pre-scan of the tensor directory throws IllegalArgumentException before any tensor is delivered, naming every offending tensor, its type (raw on-disk value for unknown types, truncated listing past 8 offenders), and the supported set:

    GGUF contains 1 tensor(s) with quantization types this loader does not support: 'blk.0.ffn_down.weight' (Q4_1). Supported types: F32, I32, F16, BF16, Q4_0, Q5_0, Q5_1, Q4_K, Q5_K, Q6_K, Q8_0. Re-quantize the model to a supported format (e.g. Q8_0, Q4_0, Q4_K or F16).

  2. Q4_0 / Q5_0 / Q5_1 had packed TensorData (fromRawBytes) and matmul kernels but were missing from the loader's when — a Q4_0 GGUF (a first-class kernel format) loaded with missing weights. They now load as packed blocks, mirroring the Q8_0/Q4_K branches.

Design notes

  • Both the pre-scan and the when anchor on a new SUPPORTED_TENSOR_TYPES companion set; the per-tensor else is now a hard IllegalStateException that names the drift if the two ever disagree.
  • Dtype-policy mismatches keep their existing skip-via-null behavior (e.g. an I32 tensor under an FP32 load) — that path is a filter consumers rely on, not an error. Only tensor types fail fast, matching the "fail before execution" rule validatePolicy/withPolicy already applies (its kdoc previously acknowledged this exact gap).
  • Behavior change: a file that previously "loaded" with skipped tensors now fails at load. CHANGELOG entry included.

Tests

  • commonTest StreamingGgufParametersLoaderFailFastTest (5 cases, runs on every target): supported set passes, Q4_1 rejected with tensor name + supported set, unknown raw type value reported verbatim, long offender lists truncated with a count, Q4_0/Q5_0/Q5_1 pinned in the supported set.
  • jvmTest: synthesized-GGUF e2e — Q4_1 file throws before any onTensorLoaded callback; multiple offenders all named; Q4_0/Q5_0/Q5_1 load as PackedBlockStorage. The existing GGUF-builder helper was parameterized to cover arbitrary tensor sets.

Verified

  • :skainet-io:skainet-io-gguf:jvmTest — 19/19 green (6 loader + 5 fail-fast + 8 policy)
  • :skainet-io:skainet-io-gguf:linuxX64Test — green (commonMain + commonTest on Kotlin/Native)
  • :skainet-io:skainet-io-core:jvmTest — green (downstream GGUF fixture loads unaffected)
  • No other in-repo consumers of the loader

Closes #919

🤖 Generated with Claude Code

…5_1 packed

StreamingGgufParametersLoader silently skipped any tensor type outside its
when-expression — a SKIP progress string was the only signal, the load
'succeeded' with missing weights, and the failure surfaced far away in the
forward pass (the load-time half of #654). The deprecated legacy loader
already hard-failed here; the streaming loader regressed that.

Two gaps hid in that else branch:

- Truly unsupported types (e.g. Q4_1): an eager pre-scan of the tensor
  directory now throws IllegalArgumentException before any tensor is
  delivered, naming every offending tensor, its type (raw value for
  unknown types), and the supported set. The per-tensor else becomes a
  hard error guarding drift from the new SUPPORTED_TENSOR_TYPES companion
  set, which both checks derive from.
- Q4_0 / Q5_0 / Q5_1 had packed TensorData (fromRawBytes) and matmul
  kernels but were absent from the loader — they now load as packed
  blocks like Q8_0/Q4_K instead of being skipped.

Dtype-policy mismatches (e.g. an I32 tensor under an FP32 load) keep the
existing skip-via-null behavior — that path is a filter, not an error.

Behavior change: files that previously 'loaded' with skipped tensors now
fail at load, matching the RFC's fail-before-execution rule already
applied by validatePolicy/withPolicy.

Verified: jvmTest (6 loader + 5 new fail-fast + 8 policy tests green),
linuxX64Test green (commonMain/commonTest on Kotlin/Native), io-core
jvmTest green (downstream GGUF fixture loads unaffected).

Closes #919

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal requested a review from aharakal August 10, 2026 05:15
aharakal
aharakal previously approved these changes Aug 10, 2026
@michalharakal
michalharakal merged commit 6ab532b into develop Aug 10, 2026
9 checks passed
@michalharakal
michalharakal deleted the fix/gguf-loader-fail-fast-919 branch August 10, 2026 08:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants