Skip to content

AOT GGUF converter for I2_S repack (eager-exec) — done; IREE leg moved out of core #1207

Description

@michalharakal

Sub-issue of #1198.

Why this, not just #1204's sidecar

The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold
start — decode + re-encode + write-to-disk, on the exact device class (constrained, e.g. a 2GB
phone) this whole effort is trying to protect. That's in tension with how this codebase already
treats cost elsewhere: PlannerProfile's own philosophy is that a cost should be visible and
paid deliberately, not quietly incurred as a surprise multiple of a weight's size (see
PlannerProfile.MOBILE_2GB's strict mode, skainet-lang-core/.../memory/plan/PlannerProfile.kt).
A silent 2x-cost first load is exactly that kind of surprise, and sits oddly next to a DSL/runtime
that otherwise insists costs be visible and chosen, not incurred quietly.

Two things already established in this thread point to a better fix:

  • NeoGPU (the vendored ternary kernel's own upstream author) never does a runtime repack at all —
    its own converter (convert_bitnet_to_gguf.py) emits the kernel-ready SEQUENTIAL layout
    offline, ahead of time.
  • The IREE-compiled leg already has a real AOT checkpoint (iree-compile) where this repack folds
    in for free — ExternalParameterRef's .irpa packager
    (skainet-compile-hlo/.../hlo/ConstantMaterialization.kt:84-90) already wants verbatim,
    kernel-ready bytes for its zero-copy path, and hits the exact same I2_S blocker the eager-exec
    loader does.

Proposed change

A shared conversion module: read any supported source format, apply whatever repack/relayout each
target dispatch key needs (I2sRepack.toSequentialPayload for ternary, Q4_K/Q5_K feed-order via
TensorView.prepack(), etc. — the same logic #1202's kernels already dispatch against), then emit
either:

  • a kernel-ready GGUF/sidecar for the eager-exec loader to consume AS_STORED — zero runtime
    transformation, real mmap via AllocationResolver's existing fileBytesAreTheBytes gate.
    skainet-io/skainet-io-gguf/.../export/GGUFWriter.kt already knows GGUF's binary format today
    (for exporting a SKaiNET-authored ComputeGraph, not yet for round-tripping an existing GGUF) —
    real infrastructure to extend, not a green field.
  • an .irpa parameter archive for the IREE leg, via skainet-io-iree-params's IrpaWriter.

Both outputs share the same repack step; only the container differs. This is the primary
mechanism whenever a build controls its own model pipeline (ships specific converted models, not
arbitrary user-supplied GGUFs) — the common production case, and the one this DSL's other export
paths (Minerva, StableHLO/IREE) already assume.

Relationship to #1204

#1204's on-device sidecar is demoted to an explicit fallback, not the default recommendation:
still needed for the case where a consumer loads a GGUF this AOT tool never saw (a user picks an
arbitrary HuggingFace file at runtime), but not the first thing to build and not what a build with
a fixed model set should reach for.

Verification

  • Round-trip test: AOT-convert a synthetic I2_S GGUF (both GROUP_128/GROUP_64 inputs), load the
    converted output through the eager-exec loader, assert decoded values match the non-converted
    reference bit-for-bit.
  • Confirm the converted output resolves to MemoryDomain.MMAP_FILE via
    AllocationResolver.explain() — no repack, no off-heap allocation, nothing to cache at load time
    at all.
  • .irpa side: confirm IrpaWriter accepts the converted, kernel-ready bytes as an
    ExternalParameterRef without re-encoding.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    sub-issueSub-issue of a tracking issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions