feat(io): I2_S GGUF import — packed BITNET_B1_58 loading with group→sequential repack (#1140) - #1165
Merged
Merged
Conversation
GGML type 36 (BitNet.cpp's I2_S) now loads as BitNetB158TensorData — the first packed ternary TensorData: 0.25 bytes per weight plus one FP32 scale, instead of the #1033 FP32 widening. Its packedView carries the exact FP32×BITNET_B1_58 dispatch key the ternary f32 kernel pack serves (#1138), so a loaded BitNet weight reaches the vendored NeoGPU LUT kernel with no further conversion. The messy part is the wire format, and it is pinned against the vendored BitNet.cpp sources rather than folklore: the payload's bit order depends on which converter wrote the file — BitNet.cpp packs QK-element blocks high bits first, and QK_I2_S is 128 on the x86 pipeline but 64 on ARM (an architecture-dependent file format); NeoGPU's converter packs sequentially, already the BITNET_B1_58 order. The scale moves too: BitNet.cpp puts one FP32 in a 32-byte trailer after the payload (w = (code-1)*scale); NeoGPU writes a companion <name>_scale tensor defined as divide-by (stored inverted, and consumed rather than delivered as a parameter). The loader takes an i2sLayout knob (GROUP_128 default | GROUP_64 | SEQUENTIAL), repacks once at load (traced as repack-i2s), and accepts either scale source in flavor-preferred order. Byte code 3 fails fast at import — I2sRepack is the only place that validates it; the kernels stay deliberately unvalidating (3 → +2 by LUT arithmetic). GGML_QUANT_SIZES sizes the I2_S payload only; the trailer is read separately from the source, so both flavors' offset math stays truthful. DequantizeTo(FP32) still widens, through the same codec decode. Tests: repack goldens against an independent reimplementation of quantize_i2_s (both QK flavors), code-3 rejection in every layout, wrong-multiple errors naming the other flavor, synthetic-GGUF end-to-end loads for all three flavors (packed data, scale resolution, companion consumption, DequantizeTo parity), and the loader/codec agreement invariant the kernel pack is defined against. Refs #1140, #1136; relates #1033 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
7 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 4 of #1136 — closes #1140. Relates to #1033 (adds a keep-packed path for I2_S without expanding TQ1/TQ2 scope).
What
GGMLQuantizationType.I2_S(36)+ payload-onlyGGML_QUANT_SIZESentry (the per-tensor scale trailer is deliberately outside the block math).BitNetB158TensorData(skainet-lang-core) — the first packed ternaryTensorData:payload + 4B LE FP32 scale, one per-tensor block, decode pinned toTernaryCodec;packedViewcarries the exactFP32×BITNET_B1_58key the [ternary-f32] Phase 2: TernaryF32GemvNative SPI + reference ViewKernel + TernaryF32KernelPack #1138 kernel pack serves.I2sRepack+I2sGgufLayout(io-gguf): group→sequential repack pinned against the vendored BitNet.cppquantize_i2_ssources — including the discovery thatQK_I2_Sis architecture-dependent (128 on the x86 pipeline, 64 on ARM), so the layout is a three-way knob:GROUP_128(default) |GROUP_64|SEQUENTIAL(NeoGPU converter, already payload-order). Byte code 3 fails fast at import — the single validation point; kernels stay unvalidating.TQ1_0,TQ2_0,BITNET_B1_58with block spec +activationhint; reference decoder and parity fixtures generated from the descriptor #1033 widening), traced asrepack-i2s; scale resolved from the BitNet.cpp 32-byte trailer (w = (code−1)·scale) or NeoGPU's companion<name>_scale(divide-by semantics → stored inverted; companion consumed, never delivered), flavor-preferred order.DequantizeTo(FP32)still widens via the same codec.Tests (all green locally)
I2sRepackTest— goldens vs an independent reimplementation of BitNet.cpp's packing for both QK flavors; code-3 rejection in every layout; wrong-multiple error names the other flavorI2sGgufLoadTest— synthetic GGUF end-to-end for all three flavors: packedBitNetB158TensorData, trailer/companion scale resolution, companion skipped as parameter,DequantizeToparity, and the loader↔codec agreement invariantBitNetB158TensorDataTest— codec round-trip, signed-code get/set, dispatch-format view, size validationio-gguf+lang-corejvm suites pass; commonMain cross-compiles (js, linuxX64)Caveat
Verified against the vendored BitNet.cpp sources, not yet against a real
microsoft/bitnet-b1.58-2B-4Tdownload (none available locally) — worth one confirmation run before relying on theGROUP_128default for that file; the fail-fast + layout knob keep a wrong guess loud.Next: #1141 benchmark scenario; transformers#336 can now receive packed BitNet weights.
🤖 Generated with Claude Code