refactor(gemma): SafeTensors loading collapses onto the engine's ShardedSafeTensorsParametersLoader (SKaiNET#1246) - #398
Merged
Conversation
…dedSafeTensorsParametersLoader (SKaiNET#1246) GemmaSafeTensorsLoader hand-rolled the per-tensor materialization the engine's single-file loader already owned — BF16/F16 widening via the GGUF module's DequantOps (a wrong dependency edge), a row-major transpose no call site ever enabled, and a PLE size guard — because the engine had no sharded ParametersLoader. SKaiNET 0.53.0 ships one (SKaiNET#1252). The loader now expresses only family policy: the HF allowlist and PLE size guard as the engine loader's tensorFilter, and the HF -> GGUF slot renaming (table-driven per layer). Every dtype decision is the engine's, driven by a DTypePolicy forwarded through ShardedSafeTensorsParametersLoader.withPolicy; GemmaNetworkLoader's SafeTensors lane validates against the same keep-native set as the GGUF lane instead of rejecting every Require(). transposeRowMajor and the DequantOps import are gone; GemmaSafeTensorsMappedPle and the PLE auto-disable semantics are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ked loader (SKaiNET#1246) A 2-shard checkpoint written with the engine's SafeTensorsWriter (BF16 projections, F32 norms, hand-written index, minimal config.json): every GGUF-named slot is populated after renaming, values round-trip exactly on bf16-representable fixtures, an unmapped INT64 vision-tower tensor is neither delivered nor allowed to trip the engine's fail-fast pre-scan, and Require(BF16) reaches the engine (native bf16 storage for projections, dense FP32 for the norms). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Brings the engine's billion-parameter export fixes (SKaiNET#1247) and the sharded SafeTensors ParametersLoader (SKaiNET#1246) into the resolved artifact set; the gemma family's loader collapse onto it follows. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This was referenced Sep 2, 2026
Merged
refactor(apertus, gemma3n): SafeTensors loading collapses onto the engine loader (SKaiNET#1246)
#401
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 2 of SKaiNET#1246 (engine side shipped in SKaiNET 0.53.0 via SKaiNET-developers/SKaiNET#1252). Includes the 0.53.0 catalog bump from #397 so it resolves standalone.
GemmaSafeTensorsLoaderno longer hand-rolls per-tensor materialization. It is now oneShardedSafeTensorsParametersLoader.withPolicy(indexPath, dtypePolicy, tensorFilter = …)run into a map by HF name, followed by the existing (now table-driven) HF → GGUF renaming:loadAndConvertTensor,transposeRowMajor(dead — every call site passedtranspose = false),loadFirstExistingIfFits, and theDequantOps/StreamingShardedSafeTensorsReaderimports — noio.ggufimport remains in this file (the module dependency stays for the GGUF lane)tensorFilter= the family allowlist plus the PLE-table size guard, so unmapped vision/audio tensors never materialize and are exempt from the engine's fail-fast pre-scan; PLE auto-disable semantics andGemmaSafeTensorsMappedPle.ktare unchangeddtypePolicy: DTypePolicy = Anyconstructor parameter (existing call sites source-compatible);GemmaNetworkLoader's SafeTensors lane now validates against{BF16, FP16}keep-native and forwards the policy, soRequire(BF16)reaches the engine instead of being rejectedTesting: new
GemmaSafeTensorsLoaderFixtureTest— a synthetic 2-shard checkpoint written with the engine'sSafeTensorsWriter(BF16 projections/embeddings, F32 norms, hand-written index + minimalconfig.json): all 23 GGUF slots present and nothing extra, tied output/embedding, exact value round-trip, an unmapped INT64 vision-tower decoy neither delivered nor rejected, andRequire(BF16)yieldingBf16DenseTensorDatafor projections with dense FP32 norms.:llm-inference:gemma:jvmTest16 suites / 47 tests green against the published 0.53.0.Follow-up: the sibling hand-rolled loaders (apertus, gemma3n, voxtral, llama
DecoderSafeTensorsLoaderLlama, llm-coreDecoderSafeTensorsLoader) take the same recipe; single-file (non-indexed) consumers useSafeTensorsParametersLoader.