Skip to content

AOT GGUF converter for I2_S — eager-exec half of #1207 - #1210

Merged
michalharakal merged 1 commit into
developfrom
feature/1207-aot-gguf-converter
Aug 29, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1207-aot-gguf-converter

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Sub-issue of #1198. Eager-exec half of #1207 (the IREE/.irpa leg is a separate follow-up, not
addressed here).

Why this instead of #1204's sidecar

The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold
start — decode + re-encode + write-to-disk — on the exact constrained device this effort exists
to protect. An AOT converter that runs once, offline, ahead of the device ever seeing the file,
avoids that entirely: a build that controls its own model pipeline gets a genuinely
zero-runtime-cost load, no sidecar, no first-load penalty. Full reasoning on #1198/#1207.

What this PR does

  • GgufTensorEntry gains rawBytes: ByteArray? as an alternative to tensor: a converter
    that already has the exact on-disk bytes for an entry (a passthrough of a quantized blob this
    writer doesn't need to interpret, or a buffer whose true size exceeds its type's formal
    GGML_QUANT_SIZES block math — an I2_S BITNET_B1_58 buffer with its trailing FP32 scale) can
    hand them over directly. GGUFWriter's expectedTensorSize/materializeTensor use
    rawBytes.size and the bytes verbatim when present, bypassing TensorFlatten (element-indexed,
    can't represent an opaque quantized blob) and GGML_QUANT_SIZES entirely. Every existing
    tensor-mode caller/test is unaffected — tensor stays the default path.
  • I2sScale.kt: extracted StreamingGgufParametersLoader's I2_S scale-resolution logic
    (resolveI2sScale, i2sCompanionScaleTensor, i2sTrailerScale, i2sTrailerScaleIsMappable)
    into shared functions, so the converter gets the identical answer the streaming loader would,
    rather than a second, potentially-drifting copy.
  • I2sAotConverter.convert(): reads an arbitrary GGUF via StreamingGGUFReader, repacks
    I2_S tensors into SEQUENTIAL+trailer order (I2sRepack — the same repack the loader already
    does, just ahead of time), drops the now-redundant companion scale tensor (keeping one would
    defeat Skip the I2_S repack copy for SEQUENTIAL-layout GGUFs, enable true mmap #1203's mmap fast path, which requires no companion to exist), and passes every other
    tensor and all KV metadata through unchanged. Produces a GgufWriteRequest for GGUFWriter.

Test plan

  • skainet-io-gguf full JVM test suite passes (including the pre-existing
    GGUFWriterRoundtripTest — unaffected by the GgufTensorEntry change)
  • New I2sAotConverterTest: converts a GROUP_128 (BitNet.cpp) source file, asserts the
    converted I2_S tensor decodes identically to the source and actually takes Skip the I2_S repack copy for SEQUENTIAL-layout GGUFs, enable true mmap #1203's mmap
    fast path afterwards (packedStorage is Storage.OffHeap under WeightResidency.MAPPED) —
    proving this is a real substitute for the on-device cache, not just the same cost moved
    earlier. A passthrough Q4_0 tensor is asserted unchanged by the round trip.

The on-device sidecar cache (#1204) pays its "convert once" cost by
doubling the first cold start on the constrained device this effort exists
to protect. An AOT converter that runs once, offline, ahead of the device
ever seeing the file, avoids that entirely -- a build that controls its
own model pipeline gets a genuinely zero-runtime-cost load, no sidecar, no
first-load penalty.

- GgufTensorEntry gains an optional `rawBytes: ByteArray?` alternative to
  `tensor`: a converter that already has the exact on-disk bytes for an
  entry (a passthrough of a quantized blob this writer doesn't need to
  interpret, or a buffer whose true size exceeds its type's formal
  GGML_QUANT_SIZES block math -- an I2_S BITNET_B1_58 buffer with its
  trailing FP32 scale) can hand them over directly. GGUFWriter's
  expectedTensorSize/materializeTensor use rawBytes.size and the bytes
  verbatim when present, bypassing TensorFlatten (which is element-indexed
  and can't represent an opaque quantized blob) and GGML_QUANT_SIZES
  entirely. Every existing tensor-mode caller/test is unaffected.
- StreamingGgufParametersLoader's I2_S scale-resolution logic
  (resolveI2sScale, i2sCompanionScaleTensor, i2sTrailerScale,
  i2sTrailerScaleIsMappable) is extracted into a shared I2sScale.kt so the
  new converter gets the identical answer the streaming loader would,
  rather than a second, potentially-drifting copy.
- I2sAotConverter.convert(): reads an arbitrary GGUF via
  StreamingGGUFReader, repacks I2_S tensors into SEQUENTIAL+trailer order
  (I2sRepack, the same repack the loader already does, just ahead of
  time), drops the now-redundant companion scale tensor (keeping one would
  defeat #1203's mmap fast path, which requires no companion to exist),
  and passes every other tensor and all KV metadata through unchanged.
  Produces a GgufWriteRequest for GGUFWriter to write out.

New I2sAotConverterTest: converts a GROUP_128 (BitNet.cpp) + companion-free
trailer source file, asserts the converted I2_S tensor decodes identically
to the source AND actually takes #1203's mmap fast path afterwards
(packedStorage is Storage.OffHeap under WeightResidency.MAPPED) -- proving
this is a real substitute for the on-device cache, not just the same cost
moved earlier. A passthrough Q4_0 tensor is asserted unchanged by the
round trip.

Scope: this is the eager-exec leg only. The IREE/.irpa leg (also named in
#1207) is not addressed here.
@michalharakal
michalharakal merged commit 1473152 into develop Aug 29, 2026
14 checks passed
@michalharakal
michalharakal deleted the feature/1207-aot-gguf-converter branch August 29, 2026 18:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant