AOT GGUF converter for I2_S — eager-exec half of #1207 - #1210
Merged
Merged
Conversation
The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold start on the constrained device this effort exists to protect. An AOT converter that runs once, offline, ahead of the device ever seeing the file, avoids that entirely -- a build that controls its own model pipeline gets a genuinely zero-runtime-cost load, no sidecar, no first-load penalty. - GgufTensorEntry gains an optional `rawBytes: ByteArray?` alternative to `tensor`: a converter that already has the exact on-disk bytes for an entry (a passthrough of a quantized blob this writer doesn't need to interpret, or a buffer whose true size exceeds its type's formal GGML_QUANT_SIZES block math -- an I2_S BITNET_B1_58 buffer with its trailing FP32 scale) can hand them over directly. GGUFWriter's expectedTensorSize/materializeTensor use rawBytes.size and the bytes verbatim when present, bypassing TensorFlatten (which is element-indexed and can't represent an opaque quantized blob) and GGML_QUANT_SIZES entirely. Every existing tensor-mode caller/test is unaffected. - StreamingGgufParametersLoader's I2_S scale-resolution logic (resolveI2sScale, i2sCompanionScaleTensor, i2sTrailerScale, i2sTrailerScaleIsMappable) is extracted into a shared I2sScale.kt so the new converter gets the identical answer the streaming loader would, rather than a second, potentially-drifting copy. - I2sAotConverter.convert(): reads an arbitrary GGUF via StreamingGGUFReader, repacks I2_S tensors into SEQUENTIAL+trailer order (I2sRepack, the same repack the loader already does, just ahead of time), drops the now-redundant companion scale tensor (keeping one would defeat #1203's mmap fast path, which requires no companion to exist), and passes every other tensor and all KV metadata through unchanged. Produces a GgufWriteRequest for GGUFWriter to write out. New I2sAotConverterTest: converts a GROUP_128 (BitNet.cpp) + companion-free trailer source file, asserts the converted I2_S tensor decodes identically to the source AND actually takes #1203's mmap fast path afterwards (packedStorage is Storage.OffHeap under WeightResidency.MAPPED) -- proving this is a real substitute for the on-device cache, not just the same cost moved earlier. A passthrough Q4_0 tensor is asserted unchanged by the round trip. Scope: this is the eager-exec leg only. The IREE/.irpa leg (also named in #1207) is not addressed here.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sub-issue of #1198. Eager-exec half of #1207 (the IREE/.irpa leg is a separate follow-up, not
addressed here).
Why this instead of #1204's sidecar
The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold
start — decode + re-encode + write-to-disk — on the exact constrained device this effort exists
to protect. An AOT converter that runs once, offline, ahead of the device ever seeing the file,
avoids that entirely: a build that controls its own model pipeline gets a genuinely
zero-runtime-cost load, no sidecar, no first-load penalty. Full reasoning on #1198/#1207.
What this PR does
GgufTensorEntrygainsrawBytes: ByteArray?as an alternative totensor: a converterthat already has the exact on-disk bytes for an entry (a passthrough of a quantized blob this
writer doesn't need to interpret, or a buffer whose true size exceeds its type's formal
GGML_QUANT_SIZESblock math — an I2_SBITNET_B1_58buffer with its trailing FP32 scale) canhand them over directly.
GGUFWriter'sexpectedTensorSize/materializeTensoruserawBytes.sizeand the bytes verbatim when present, bypassingTensorFlatten(element-indexed,can't represent an opaque quantized blob) and
GGML_QUANT_SIZESentirely. Every existingtensor-mode caller/test is unaffected —
tensorstays the default path.I2sScale.kt: extractedStreamingGgufParametersLoader's I2_S scale-resolution logic(
resolveI2sScale,i2sCompanionScaleTensor,i2sTrailerScale,i2sTrailerScaleIsMappable)into shared functions, so the converter gets the identical answer the streaming loader would,
rather than a second, potentially-drifting copy.
I2sAotConverter.convert(): reads an arbitrary GGUF viaStreamingGGUFReader, repacksI2_S tensors into
SEQUENTIAL+trailer order (I2sRepack— the same repack the loader alreadydoes, just ahead of time), drops the now-redundant companion scale tensor (keeping one would
defeat Skip the I2_S repack copy for SEQUENTIAL-layout GGUFs, enable true mmap #1203's mmap fast path, which requires no companion to exist), and passes every other
tensor and all KV metadata through unchanged. Produces a
GgufWriteRequestforGGUFWriter.Test plan
skainet-io-gguffull JVM test suite passes (including the pre-existingGGUFWriterRoundtripTest— unaffected by theGgufTensorEntrychange)I2sAotConverterTest: converts aGROUP_128(BitNet.cpp) source file, asserts theconverted I2_S tensor decodes identically to the source and actually takes Skip the I2_S repack copy for SEQUENTIAL-layout GGUFs, enable true mmap #1203's mmap
fast path afterwards (
packedStorage is Storage.OffHeapunderWeightResidency.MAPPED) —proving this is a real substitute for the on-device cache, not just the same cost moved
earlier. A passthrough Q4_0 tensor is asserted unchanged by the round trip.