Skip to content

refactor(llm-core): DecoderSafeTensorsLoader collapses onto the engine loader (SKaiNET#1246) - #400

Merged
michalharakal merged 3 commits into
developfrom
feat/1246-decoder-safetensors-loaders
Sep 2, 2026
Merged

michalharakal merged 3 commits into
developfrom
feat/1246-decoder-safetensors-loaders

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

#1246 Phase 2, step 5 (part 1 of 2) — follows #398's recipe for the shared decoder loader.

  • DecoderSafeTensorsLoader.loadToMap does a header-only pass, then for every all-float (F32/F16/BF16) file runs the engine's SafeTensorsParametersLoader.withPolicy(provider, dtypePolicy) — the engine owns reading and the BF16/F16 widen-vs-keep-native decision. Family-side adapt() keeps only what the engine doesn't know: HF→GGUF renaming (unmapped tensors dropped), [1, dim] norm normalization (re-wrap, no copy), tied embeddings, and the input-major relayout of keep-native rank-2 matmul weights (engine #888). Public signature unchanged (loadToMap stays non-suspend for KLlamaJava).
  • Retained as loadLegacy: the SKaiNET-specific Q4 + .qb companion export and files with non-float/F64 tensors — the engine's single-file loader has no tensorFilter and would throw mid-load (filed as SafeTensors: tensorFilter on the single-file SafeTensorsParametersLoader (parity with the sharded loader) SKaiNET#1256; with it the legacy path shrinks to Q4-only).
  • DecoderSafeTensorsLoaderLlama had no materialization of its own (a 10-line LlamaWeightMapper.map(loadToMap(...))), so it gains only a fixture test pinning the engine loader → renaming → mapper contract.

Testing (published 0.53.0): new DecoderSafeTensorsLoaderFixtureTest (3: slots/norms/tied/unmapped-dropped/exact values; Require(BF16) → NarrowFloatInputMajorTensorData with Bf16Codec for projections; legacy Q4+.qb still dequantizes) and DecoderSafeTensorsLoaderLlamaFixtureTest (2). :llm-core:jvmTest 130/0, :llm-inference:llama:jvmTest 35/0, kllama + skainet-decode-core compile. :llm-core:apiDump — no diff.

michalharakal and others added 3 commits September 2, 2026 14:56
…der (SKaiNET#1246)

DecoderSafeTensorsLoader.loadToMap rides SafeTensorsParametersLoader.withPolicy
for every all-float checkpoint: the engine owns reading and the BF16/F16
widen-or-keep-native decision for the DTypePolicy; the family keeps only what
the engine does not know — HF -> GGUF renaming, [1, dim] norm normalization,
tied embeddings, and the input-major relayout of keep-native matmul weights
(engine #888). The public surface is unchanged (loadToMap stays a plain
function for the Java consumers: the engine load is suspend by contract but
never suspends, so it is driven to completion without a coroutine runtime).

The legacy Q4 + .qb export, non-float tensors and non-FP32 witnesses cannot
ride the engine's single-file loader (no tensorFilter, no Q4 dtype); those
files keep the pre-#1246 reader verbatim, scoped as loadLegacy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nsors load into LlamaRuntimeWeights (SKaiNET#1246)

DecoderSafeTensorsLoaderLlama is a 10-line mapping extension with no
materialization of its own; this pins the engine loader -> family renaming ->
LlamaWeightMapper contract end to end for both policy arms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit a0c9309 into develop Sep 2, 2026
2 checks passed
@michalharakal
michalharakal deleted the feat/1246-decoder-safetensors-loaders branch September 2, 2026 13:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant