Sub-issue of #1198.
Why this, not just #1204's sidecar
The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold
start — decode + re-encode + write-to-disk, on the exact device class (constrained, e.g. a 2GB
phone) this whole effort is trying to protect. That's in tension with how this codebase already
treats cost elsewhere: PlannerProfile's own philosophy is that a cost should be visible and
paid deliberately, not quietly incurred as a surprise multiple of a weight's size (see
PlannerProfile.MOBILE_2GB's strict mode, skainet-lang-core/.../memory/plan/PlannerProfile.kt).
A silent 2x-cost first load is exactly that kind of surprise, and sits oddly next to a DSL/runtime
that otherwise insists costs be visible and chosen, not incurred quietly.
Two things already established in this thread point to a better fix:
- NeoGPU (the vendored ternary kernel's own upstream author) never does a runtime repack at all —
its own converter (convert_bitnet_to_gguf.py) emits the kernel-ready SEQUENTIAL layout
offline, ahead of time.
- The IREE-compiled leg already has a real AOT checkpoint (
iree-compile) where this repack folds
in for free — ExternalParameterRef's .irpa packager
(skainet-compile-hlo/.../hlo/ConstantMaterialization.kt:84-90) already wants verbatim,
kernel-ready bytes for its zero-copy path, and hits the exact same I2_S blocker the eager-exec
loader does.
Proposed change
A shared conversion module: read any supported source format, apply whatever repack/relayout each
target dispatch key needs (I2sRepack.toSequentialPayload for ternary, Q4_K/Q5_K feed-order via
TensorView.prepack(), etc. — the same logic #1202's kernels already dispatch against), then emit
either:
- a kernel-ready GGUF/sidecar for the eager-exec loader to consume
AS_STORED — zero runtime
transformation, real mmap via AllocationResolver's existing fileBytesAreTheBytes gate.
skainet-io/skainet-io-gguf/.../export/GGUFWriter.kt already knows GGUF's binary format today
(for exporting a SKaiNET-authored ComputeGraph, not yet for round-tripping an existing GGUF) —
real infrastructure to extend, not a green field.
- an
.irpa parameter archive for the IREE leg, via skainet-io-iree-params's IrpaWriter.
Both outputs share the same repack step; only the container differs. This is the primary
mechanism whenever a build controls its own model pipeline (ships specific converted models, not
arbitrary user-supplied GGUFs) — the common production case, and the one this DSL's other export
paths (Minerva, StableHLO/IREE) already assume.
Relationship to #1204
#1204's on-device sidecar is demoted to an explicit fallback, not the default recommendation:
still needed for the case where a consumer loads a GGUF this AOT tool never saw (a user picks an
arbitrary HuggingFace file at runtime), but not the first thing to build and not what a build with
a fixed model set should reach for.
Verification
- Round-trip test: AOT-convert a synthetic I2_S GGUF (both
GROUP_128/GROUP_64 inputs), load the
converted output through the eager-exec loader, assert decoded values match the non-converted
reference bit-for-bit.
- Confirm the converted output resolves to
MemoryDomain.MMAP_FILE via
AllocationResolver.explain() — no repack, no off-heap allocation, nothing to cache at load time
at all.
.irpa side: confirm IrpaWriter accepts the converted, kernel-ready bytes as an
ExternalParameterRef without re-encoding.
Sub-issue of #1198.
Why this, not just #1204's sidecar
The on-device sidecar cache (#1204) pays its "convert once" cost by doubling the first cold
start — decode + re-encode + write-to-disk, on the exact device class (constrained, e.g. a 2GB
phone) this whole effort is trying to protect. That's in tension with how this codebase already
treats cost elsewhere:
PlannerProfile's own philosophy is that a cost should be visible andpaid deliberately, not quietly incurred as a surprise multiple of a weight's size (see
PlannerProfile.MOBILE_2GB'sstrictmode,skainet-lang-core/.../memory/plan/PlannerProfile.kt).A silent 2x-cost first load is exactly that kind of surprise, and sits oddly next to a DSL/runtime
that otherwise insists costs be visible and chosen, not incurred quietly.
Two things already established in this thread point to a better fix:
its own converter (
convert_bitnet_to_gguf.py) emits the kernel-readySEQUENTIALlayoutoffline, ahead of time.
iree-compile) where this repack foldsin for free —
ExternalParameterRef's.irpapackager(
skainet-compile-hlo/.../hlo/ConstantMaterialization.kt:84-90) already wants verbatim,kernel-ready bytes for its zero-copy path, and hits the exact same I2_S blocker the eager-exec
loader does.
Proposed change
A shared conversion module: read any supported source format, apply whatever repack/relayout each
target dispatch key needs (
I2sRepack.toSequentialPayloadfor ternary, Q4_K/Q5_K feed-order viaTensorView.prepack(), etc. — the same logic #1202's kernels already dispatch against), then emiteither:
AS_STORED— zero runtimetransformation, real mmap via
AllocationResolver's existingfileBytesAreTheBytesgate.skainet-io/skainet-io-gguf/.../export/GGUFWriter.ktalready knows GGUF's binary format today(for exporting a SKaiNET-authored
ComputeGraph, not yet for round-tripping an existing GGUF) —real infrastructure to extend, not a green field.
.irpaparameter archive for the IREE leg, viaskainet-io-iree-params'sIrpaWriter.Both outputs share the same repack step; only the container differs. This is the primary
mechanism whenever a build controls its own model pipeline (ships specific converted models, not
arbitrary user-supplied GGUFs) — the common production case, and the one this DSL's other export
paths (Minerva, StableHLO/IREE) already assume.
Relationship to #1204
#1204's on-device sidecar is demoted to an explicit fallback, not the default recommendation:
still needed for the case where a consumer loads a GGUF this AOT tool never saw (a user picks an
arbitrary HuggingFace file at runtime), but not the first thing to build and not what a build with
a fixed model set should reach for.
Verification
GROUP_128/GROUP_64inputs), load theconverted output through the eager-exec loader, assert decoded values match the non-converted
reference bit-for-bit.
MemoryDomain.MMAP_FILEviaAllocationResolver.explain()— no repack, no off-heap allocation, nothing to cache at load timeat all.
.irpaside: confirmIrpaWriteraccepts the converted, kernel-ready bytes as anExternalParameterRefwithout re-encoding.