feat(io-safetensors): sharded-index ParametersLoader riding openFromIndex (#1246) - #1252
Merged
Merged
Conversation
…d materializer (#1246) Extract the per-tensor dtype dispatch, byte/dequant helpers, and DTypePolicy mappers from SafeTensorsParametersLoader into an internal SafeTensorsMaterializer (pure refactor; primitive-typed signature so it serves both StreamingSafeTensorInfo and ShardedTensorInfo), then add ShardedSafeTensorsParametersLoader consuming model.safetensors.index.json via StreamingShardedSafeTensorsReader.openFromIndex. The sharded loader adds a fail-fast dtype pre-scan (one aggregated error before any tensor is delivered, mirroring the GGUF loader's #919 contract) and a tensorFilter hook so family-side skip policy (size guards, name allowlists) stays out of the engine while all dtype/policy handling stays in it. withPolicy has signature parity with the single-file factory. Tests: commonTest policy-routing + pre-scan suite; jvmTest 2-shard fixture (cross-shard name-sorted delivery, KEEP_NATIVE vs DEQUANT arms, tensorFilter, IncompleteShard vs allowPartial, fail-fast-before-first- delivery) and a 1-shard-index vs single-file parity test guarding the extraction refactor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Sep 2, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1246.
SafeTensorsParametersLoaderis single-file only while the sharded machinery (StreamingShardedSafeTensorsReader.openFromIndex,SafeTensorsIndexParser) already lives engine-side — so transformers families with sharded HF checkpoints hand-roll per-tensor materialization (GemmaSafeTensorsLoaderet al.), duplicating the policy surface this abstraction owns.SafeTensorsMaterializer(internal, commonMain): the dtype-dispatch materialization, byte/dequant helpers, andDTypePolicymappers extracted from the single-file loader — primitive-typed so it serves bothStreamingSafeTensorInfoandShardedTensorInfo. The single-file loader delegates; its public surface is unchanged (pure refactor, existing tests pass untouched).ShardedSafeTensorsParametersLoader(indexPath, onProgress, bf16Policy, fp16Policy, allowPartial, tensorFilter): ridesopenFromIndex, fail-fast aggregated dtype pre-scan before any tensor is delivered (mirrors the GGUF loader's StreamingGgufParametersLoader silently skips tensors it can't load (Q4_1, but also Q4_0/Q5_0/Q5_1) — load "succeeds" with missing weights, crash comes later #919 behavior),withPolicycompanion with signature parity.tensorFilterkeeps skip decisions (size guards, name allowlists) family-side while all dtype/policy handling stays engine-side (Bump org.jetbrains.kotlinx.kover from 0.9.4 to 0.9.5 #346's "no per-family quant code").Testing: 145 module tests green — commonTest policy-mapper parity, and a jvmTest suite over a genuine 2-shard fixture (cross-shard delivery/values/order, KEEP_NATIVE vs DEQUANT storage types,
tensorFilter, missing-shardIncompleteShardvsallowPartial, fail-fast with zero deliveries) plus a 1-shard-index vs single-file parity test guarding the refactor.Follow-up (transformers, after this ships in a release): collapse
GemmaSafeTensorsLoaderand the sibling hand-rolled loaders onto this, and flipGemmaNetworkLoader'skeepNative = emptySet()tosetOf(BF16, FP16).