feat(io): NameMap — GGUF/HF tensor names ↔ TensorId for Llama, Qwen2, Gemma-3; unmapped reported (SKEEP-003 P1, S0.6) - #1055
Merged
Conversation
Contributor
Author
|
Local gate Targeted: io-core |
… Gemma-3; unmapped names reported (SKEEP-003 P1) Milestone M0 (#1001), PRD M0-F4 / M0-A2 (all tensors of Llama-3.2-1B, Qwen2.5-0.5B, Gemma-3-1B map to TensorIds, zero unmapped). - io-core sk.ainet.io.weights.NameMap: bidirectional checkpoint name <-> TensorId, unmapped(names) (never dropped), toTensorIds, and asWeightNameResolver() adapter to the legacy (modulePath, paramName) resolver. RoleTableNameMap: role tables for top-level and per-layer tensors with a {N} layer-prefix pattern; TransformerNameMaps.Gguf (llama, qwen2, gemma3; forArchitecture) and .Hf (llama, qwen2, gemma3). Canonical ids are family-neutral, GGUF-role based (model.layers[N].attn.q_proj.weight, model.layers[N].post_ffw_norm.weight, model.embed_tokens.weight, model.lm_head.weight, model.rope_freqs.weight); HF norms are mapped per family (Gemma-3's post_attention_layernorm is a true post-attention norm, Llama/Qwen2's is the pre-FFN ffn_norm). - io-gguf: StreamingGGUFReader.nameMap() (from general.architecture), tensorIds(map), NameMap.asTensorNameMapper() adapter. - Tests: NameMapTest (full synthetic tensor lists of the three reference GGUFs incl. Qwen2 q/k/v biases and Gemma-3's four norms + q/k norms, round trips, family-neutral ids, HF norm mapping, unmapped reporting, resolver adapter); GgufNameMapFixtureTest (fixture-gated real files, -Dskainet.test.fixturesDir, [skip] when absent). Closes #1011 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
michalharakal
force-pushed
the
feature/1010-tensor-id
branch
from
August 22, 2026 21:08
e821ffb to
6ca4ed9
Compare
michalharakal
force-pushed
the
feature/1011-gguf-namemap
branch
from
August 22, 2026 21:08
a9d4821 to
fae03eb
Compare
Contributor
Author
|
Rebased onto the updated base after #1053 merged into |
This was referenced Aug 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S0.6 (milestone M0 #1001, PRD M0-F4 / M0-A2):
NameMap— checkpoint tensor names ↔TensorId, bidirectional, for the three reference families, with unmapped names reported, never dropped.sk.ainet.io.weights.NameMap:toTensorId(name),toCheckpointName(id),unmapped(names),toTensorIds(names),asWeightNameResolver()(adapter to the legacy(modulePath, paramName)resolver).RoleTableNameMap: role tables for top-level and per-layer tensors with a{N}layer-prefix pattern.TransformerNameMaps.Gguf(llama,qwen2,gemma3,forArchitecture(general.architecture)) and.Hf(llama,qwen2,gemma3).model.layers[N].attn.{q,k,v,o}_proj.{weight,bias},model.layers[N].attn.{q,k}_norm.weight,model.layers[N].mlp.{gate,up,down}_proj.weight,model.layers[N].{attn_norm,ffn_norm,post_attention_norm,post_ffw_norm}.weight,model.{embed_tokens,norm,lm_head,rope_freqs}.weight. HF norms are mapped per family (Gemma-3'spost_attention_layernormis a true post-attention norm; Llama/Qwen2's is the pre-FFNffn_norm).StreamingGGUFReader.nameMap()(fromgeneral.architecture),tensorIds(map),NameMap.asTensorNameMapper()(adapter to the legacy role interface).NameMapTest— the full tensor lists of Llama-3.2-1B (16 layers,rope_freqs), Qwen2.5-0.5B (24 layers, q/k/v biases), Gemma-3-1B (26 layers, q/k norms, four norms) map with zero unmapped and round-trip; family-neutral ids; HF norm mapping; unmapped reporting; resolver adapter.GgufNameMapFixtureTest— the same on real files, fixture-gated (-Dskainet.test.fixturesDir; the Qwen2.5-0.5B Q8_0 GGUF fromdownloadQwenTokenizerFixturesis picked up automatically; gated Llama/Gemma files run when dropped in).Stacked on #1054 (S0.5 —
TensorId); retarget todevelopafter it merges.Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25); results in the first comment. Targeted: io-coresk.ainet.io.weights.*7/7, io-gguf fixture test 3/3 (skips when files absent).Closes #1011
🤖 Generated with Claude Code