refactor(llm-core): DecoderSafeTensorsLoader collapses onto the engine loader (SKaiNET#1246) - #400
Merged
Conversation
…der (SKaiNET#1246) DecoderSafeTensorsLoader.loadToMap rides SafeTensorsParametersLoader.withPolicy for every all-float checkpoint: the engine owns reading and the BF16/F16 widen-or-keep-native decision for the DTypePolicy; the family keeps only what the engine does not know — HF -> GGUF renaming, [1, dim] norm normalization, tied embeddings, and the input-major relayout of keep-native matmul weights (engine #888). The public surface is unchanged (loadToMap stays a plain function for the Java consumers: the engine load is suspend by contract but never suspends, so it is driven to completion without a coroutine runtime). The legacy Q4 + .qb export, non-float tensors and non-FP32 witnesses cannot ride the engine's single-file loader (no tensorFilter, no Q4 dtype); those files keep the pre-#1246 reader verbatim, scoped as loadLegacy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nsors load into LlamaRuntimeWeights (SKaiNET#1246) DecoderSafeTensorsLoaderLlama is a 10-line mapping extension with no materialization of its own; this pins the engine loader -> family renaming -> LlamaWeightMapper contract end to end for both policy arms. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…safetensors-loaders
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#1246 Phase 2, step 5 (part 1 of 2) — follows #398's recipe for the shared decoder loader.
DecoderSafeTensorsLoader.loadToMapdoes a header-only pass, then for every all-float (F32/F16/BF16) file runs the engine'sSafeTensorsParametersLoader.withPolicy(provider, dtypePolicy)— the engine owns reading and the BF16/F16 widen-vs-keep-native decision. Family-sideadapt()keeps only what the engine doesn't know: HF→GGUF renaming (unmapped tensors dropped),[1, dim]norm normalization (re-wrap, no copy), tied embeddings, and the input-major relayout of keep-native rank-2 matmul weights (engine #888). Public signature unchanged (loadToMapstays non-suspend forKLlamaJava).loadLegacy: the SKaiNET-specificQ4+.qbcompanion export and files with non-float/F64 tensors — the engine's single-file loader has notensorFilterand would throw mid-load (filed as SafeTensors: tensorFilter on the single-file SafeTensorsParametersLoader (parity with the sharded loader) SKaiNET#1256; with it the legacy path shrinks to Q4-only).DecoderSafeTensorsLoaderLlamahad no materialization of its own (a 10-lineLlamaWeightMapper.map(loadToMap(...))), so it gains only a fixture test pinning the engine loader → renaming → mapper contract.Testing (published 0.53.0): new
DecoderSafeTensorsLoaderFixtureTest(3: slots/norms/tied/unmapped-dropped/exact values;Require(BF16)→NarrowFloatInputMajorTensorDatawithBf16Codecfor projections; legacy Q4+.qb still dequantizes) andDecoderSafeTensorsLoaderLlamaFixtureTest(2).:llm-core:jvmTest130/0,:llm-inference:llama:jvmTest35/0, kllama + skainet-decode-core compile.:llm-core:apiDump— no diff.