Skip to content

feat(io): I2_S GGUF import — packed BITNET_B1_58 loading with group→sequential repack (#1140) - #1165

Merged
michalharakal merged 1 commit into
developfrom
feature/1140-i2s-gguf-import
Aug 26, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1140-i2s-gguf-import

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Phase 4 of #1136 — closes #1140. Relates to #1033 (adds a keep-packed path for I2_S without expanding TQ1/TQ2 scope).

What

  • GGMLQuantizationType.I2_S(36) + payload-only GGML_QUANT_SIZES entry (the per-tensor scale trailer is deliberately outside the block math).
  • BitNetB158TensorData (skainet-lang-core) — the first packed ternary TensorData: payload + 4B LE FP32 scale, one per-tensor block, decode pinned to TernaryCodec; packedView carries the exact FP32×BITNET_B1_58 key the [ternary-f32] Phase 2: TernaryF32GemvNative SPI + reference ViewKernel + TernaryF32KernelPack #1138 kernel pack serves.
  • I2sRepack + I2sGgufLayout (io-gguf): group→sequential repack pinned against the vendored BitNet.cpp quantize_i2_s sources — including the discovery that QK_I2_S is architecture-dependent (128 on the x86 pipeline, 64 on ARM), so the layout is a three-way knob: GROUP_128 (default) | GROUP_64 | SEQUENTIAL (NeoGPU converter, already payload-order). Byte code 3 fails fast at import — the single validation point; kernels stay unvalidating.
  • Loader branch: I2_S loads packed (0.25 B/weight — first ternary type without the [S2.7] P6: ternary encodings TQ1_0, TQ2_0, BITNET_B1_58 with block spec + activation hint; reference decoder and parity fixtures generated from the descriptor #1033 widening), traced as repack-i2s; scale resolved from the BitNet.cpp 32-byte trailer (w = (code−1)·scale) or NeoGPU's companion <name>_scale (divide-by semantics → stored inverted; companion consumed, never delivered), flavor-preferred order. DequantizeTo(FP32) still widens via the same codec.

Tests (all green locally)

  • I2sRepackTest — goldens vs an independent reimplementation of BitNet.cpp's packing for both QK flavors; code-3 rejection in every layout; wrong-multiple error names the other flavor
  • I2sGgufLoadTest — synthetic GGUF end-to-end for all three flavors: packed BitNetB158TensorData, trailer/companion scale resolution, companion skipped as parameter, DequantizeTo parity, and the loader↔codec agreement invariant
  • BitNetB158TensorDataTest — codec round-trip, signed-code get/set, dispatch-format view, size validation
  • Full io-gguf + lang-core jvm suites pass; commonMain cross-compiles (js, linuxX64)

Caveat

Verified against the vendored BitNet.cpp sources, not yet against a real microsoft/bitnet-b1.58-2B-4T download (none available locally) — worth one confirmation run before relying on the GROUP_128 default for that file; the fail-fast + layout knob keep a wrong guess loud.

Next: #1141 benchmark scenario; transformers#336 can now receive packed BitNet weights.

🤖 Generated with Claude Code

GGML type 36 (BitNet.cpp's I2_S) now loads as BitNetB158TensorData — the
first packed ternary TensorData: 0.25 bytes per weight plus one FP32 scale,
instead of the #1033 FP32 widening. Its packedView carries the exact
FP32×BITNET_B1_58 dispatch key the ternary f32 kernel pack serves (#1138),
so a loaded BitNet weight reaches the vendored NeoGPU LUT kernel with no
further conversion.

The messy part is the wire format, and it is pinned against the vendored
BitNet.cpp sources rather than folklore: the payload's bit order depends on
which converter wrote the file — BitNet.cpp packs QK-element blocks high
bits first, and QK_I2_S is 128 on the x86 pipeline but 64 on ARM (an
architecture-dependent file format); NeoGPU's converter packs sequentially,
already the BITNET_B1_58 order. The scale moves too: BitNet.cpp puts one
FP32 in a 32-byte trailer after the payload (w = (code-1)*scale); NeoGPU
writes a companion <name>_scale tensor defined as divide-by (stored
inverted, and consumed rather than delivered as a parameter). The loader
takes an i2sLayout knob (GROUP_128 default | GROUP_64 | SEQUENTIAL),
repacks once at load (traced as repack-i2s), and accepts either scale
source in flavor-preferred order. Byte code 3 fails fast at import —
I2sRepack is the only place that validates it; the kernels stay
deliberately unvalidating (3 → +2 by LUT arithmetic).

GGML_QUANT_SIZES sizes the I2_S payload only; the trailer is read
separately from the source, so both flavors' offset math stays truthful.
DequantizeTo(FP32) still widens, through the same codec decode.

Tests: repack goldens against an independent reimplementation of
quantize_i2_s (both QK flavors), code-3 rejection in every layout,
wrong-multiple errors naming the other flavor, synthetic-GGUF end-to-end
loads for all three flavors (packed data, scale resolution, companion
consumption, DequantizeTo parity), and the loader/codec agreement invariant
the kernel pack is defined against.

Refs #1140, #1136; relates #1033

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1165 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit b84b00c into develop Aug 26, 2026
15 of 16 checks passed
@michalharakal
michalharakal deleted the feature/1140-i2s-gguf-import branch August 26, 2026 11:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ternary-f32] Phase 4: GGUF I2_S (type 36) import with group→sequential repack, keep-packed BITNET_B1_58

1 participant