You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Each model family module grew its own structure organically. With the 0.49.0 memory-model adoption (#338–#344) rewriting every family's loading path, now is the moment to fix one common scheme so every family looks the same and adding a new one is mechanical.
De-facto convention (Apertus/Voxtral/Gemma follow it most closely)
bare descriptive names in the family module (ApertusXIELU.kt, AltUp.kt); shared ones live in transformer-core / llm-core
—
Runtime facade
<F>Ingestion in llm-runtime/k<f>
—
Weight materialization is delegated to the engine (StreamingGgufParametersLoader / WeightForm from sk.ainet.core:skainet-io-gguf); no per-family quant/packing/converter code (*QuantLayout, *PackedWeights, *MemSegConverter, QuantizedTensorFactory are legacy and retire with the 0.49.0 migration).
Known drift to converge
llama/qwen/smollm2: shared DecoderGgufWeightLoader instead of family prefix; QwenGgufWeightSource naming; three different tensor-name patterns (LlamaGgufTensorNames vs Gemma3nGgufTensorNames vs ApertusTensorNames vs VoxtralTensorNames).
gemma: versioned + unversioned prefixes mixed in one module (Gemma4WeightLoader, Gemma3nWeightLoader, GemmaNetworkLoader).
Motivation
Each model family module grew its own structure organically. With the 0.49.0 memory-model adoption (#338–#344) rewriting every family's loading path, now is the moment to fix one common scheme so every family looks the same and adding a new one is mechanical.
De-facto convention (Apertus/Voxtral/Gemma follow it most closely)
For a family
<F>(modulellm-inference/<f>):<f>Network()builder fn<F>NetworkDef.kt<F>NetworkLoader(fromGguf/fromSafeTensors/fromWeights)<F>NetworkLoader.kt<F>WeightLoader(fromSource/fromRandomAccess,loadToMap)<F>WeightLoader.kt<F>Weights,<F>RuntimeWeights,<F>TensorNames<F>RuntimeWeights.kt<F>ConfigParser<F>ConfigParser.ktApertusXIELU.kt,AltUp.kt); shared ones live intransformer-core/llm-core<F>Ingestioninllm-runtime/k<f>Weight materialization is delegated to the engine (
StreamingGgufParametersLoader/WeightFormfromsk.ainet.core:skainet-io-gguf); no per-family quant/packing/converter code (*QuantLayout,*PackedWeights,*MemSegConverter,QuantizedTensorFactoryare legacy and retire with the 0.49.0 migration).Known drift to converge
DecoderGgufWeightLoaderinstead of family prefix;QwenGgufWeightSourcenaming; three different tensor-name patterns (LlamaGgufTensorNamesvsGemma3nGgufTensorNamesvsApertusTensorNamesvsVoxtralTensorNames).Gemma4WeightLoader,Gemma3nWeightLoader,GemmaNetworkLoader).QuantizedTensorFactory/LlamaQuantLayout, gemmaGemmaQuantLayout/BlockQuantPackinguse, apertusApertusMemSegConverter— the latter already removed in 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339).Plan
This issue is the single reference for the convention; deviations should be justified in PRs against it.