Skip to content

Common structure and naming convention for model family modules #346

Description

@michalharakal

Motivation

Each model family module grew its own structure organically. With the 0.49.0 memory-model adoption (#338–#344) rewriting every family's loading path, now is the moment to fix one common scheme so every family looks the same and adding a new one is mechanical.

De-facto convention (Apertus/Voxtral/Gemma follow it most closely)

For a family <F> (module llm-inference/<f>):

Concern Name File
DSL network definition <f>Network() builder fn <F>NetworkDef.kt
End-to-end module loader <F>NetworkLoader (fromGguf / fromSafeTensors / fromWeights) <F>NetworkLoader.kt
GGUF weight materialization <F>WeightLoader (fromSource / fromRandomAccess, loadToMap) <F>WeightLoader.kt
Weight containers + tensor names <F>Weights, <F>RuntimeWeights, <F>TensorNames <F>RuntimeWeights.kt
HF config parsing <F>ConfigParser <F>ConfigParser.kt
Family-specific layers/ops bare descriptive names in the family module (ApertusXIELU.kt, AltUp.kt); shared ones live in transformer-core / llm-core —
Runtime facade <F>Ingestion in llm-runtime/k<f> —

Weight materialization is delegated to the engine (StreamingGgufParametersLoader / WeightForm from sk.ainet.core:skainet-io-gguf); no per-family quant/packing/converter code (*QuantLayout, *PackedWeights, *MemSegConverter, QuantizedTensorFactory are legacy and retire with the 0.49.0 migration).

Known drift to converge

  • llama/qwen/smollm2: shared DecoderGgufWeightLoader instead of family prefix; QwenGgufWeightSource naming; three different tensor-name patterns (LlamaGgufTensorNames vs Gemma3nGgufTensorNames vs ApertusTensorNames vs VoxtralTensorNames).
  • gemma: versioned + unversioned prefixes mixed in one module (Gemma4WeightLoader, Gemma3nWeightLoader, GemmaNetworkLoader).
  • Per-family quant machinery (llama QuantizedTensorFactory/LlamaQuantLayout, gemma GemmaQuantLayout/BlockQuantPacking use, apertus ApertusMemSegConverter — the latter already removed in 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339).

Plan

  1. 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339 (Apertus) establishes the reference shape — done as part of the B1 migration.
  2. 0.49.0 adoption B2: llama/qwen/smollm2 onto the engine loader (retire QuantizedTensorFactory) #340/0.49.0 adoption B3: gemma4/gemma3n onto the engine loader (retire BlockQuantPacking + capability gate) #341 align llama/qwen/smollm2 and gemma to this skeleton while migrating them onto the engine loader (rename where cheap; note deliberate exceptions here when a rename is not worth the churn).
  3. The BitNet module ([bitnet] T1: BitNet architecture — ModelFamily, network def, packed-weight loader #336) is the first greenfield instantiation of the template.
  4. 0.49.0 adoption E: docs sweep + load diagnostics (explainPlacements, TraceSession.identify) #344 documents the resulting scheme as the "how to add a model family" guide.

This issue is the single reference for the convention; deviations should be justified in PRs against it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions