Follow-up to #1246 / #1252, found while collapsing the SKaiNET-transformers family loaders onto the engine loaders.
ShardedSafeTensorsParametersLoader (0.53.0) takes a tensorFilter: ((ShardedTensorInfo) -> Boolean)? so a family can keep its name allowlist / size guards family-side while the engine owns every dtype decision. The single-file SafeTensorsParametersLoader has no equivalent: it materializes every tensor in the file and its per-arm dtype require throws on the first tensor the requested dtype can't accept.
Consequences downstream (SKaiNET-transformers):
Proposal (~10 lines, mirrors the sharded loader): tensorFilter: ((StreamingSafeTensorInfo) -> Boolean)? = null on SafeTensorsParametersLoader and its withPolicy companion, applied before delivery and before any fail-fast pre-scan. With it, Voxtral becomes engine loader for BF16/F16/F32 + a family-side Q4 path over the engine reader, and llm-core's legacy path shrinks to Q4-only. A SafeTensorsParametersLoader fail-fast pre-scan (the sharded loader has one, #919-style) would be the natural companion.
Follow-up to #1246 / #1252, found while collapsing the SKaiNET-transformers family loaders onto the engine loaders.
ShardedSafeTensorsParametersLoader(0.53.0) takes atensorFilter: ((ShardedTensorInfo) -> Boolean)?so a family can keep its name allowlist / size guards family-side while the engine owns every dtype decision. The single-fileSafeTensorsParametersLoaderhas no equivalent: it materializes every tensor in the file and its per-arm dtyperequirethrows on the first tensor the requesteddtypecan't accept.Consequences downstream (SKaiNET-transformers):
VoxtralSafeTensorsLoader(single file, must skip unmapped tensors and a customQUANT4+.qb-scales format) cannot collapse at all.DecoderSafeTensorsLoader(llm-core) collapsed only for all-float files; files with non-float / F64 / Q4 tensors keep the pre-SafeTensors: a sharded-index ParametersLoader (the single-file loader can't consume model.safetensors.index.json) #1246 reader verbatim as a legacy path.ApertusSingleSafeTensorsLoader(warn-and-skip on unknown dtypes) stays hand-rolled.Proposal (~10 lines, mirrors the sharded loader):
tensorFilter: ((StreamingSafeTensorInfo) -> Boolean)? = nullonSafeTensorsParametersLoaderand itswithPolicycompanion, applied before delivery and before any fail-fast pre-scan. With it, Voxtral becomes engine loader for BF16/F16/F32 + a family-side Q4 path over the engine reader, and llm-core's legacy path shrinks to Q4-only. ASafeTensorsParametersLoaderfail-fast pre-scan (the sharded loader has one, #919-style) would be the natural companion.