add GLM-5.3-Flash (GLM5-Next) support - #27773
Conversation
106ece6 to
9370c82
Compare
|
Hey @timkhronos great work on the PR! A few requests if possible:
Tagging @ngxson for visibility as well. I re-checked and if (1) + (2) is applied, the quants we uploaded work fine (+ the small shard-1 rewrite) and KLD / PPL are correct under this PR. Seems like a simple alias isn't possible actually :( It breaks the quants made with this PR |
|
Hmmm https://github.com/timkhronos/llama.cpp/pull/9/changes would alias the tensors but it looks a bit problematic hmmm |
|
Throwing up some performance numbers here from the lower end of consumer hardware (128GB DDR5 + 24GB VRAM (4090)). avar6 has some freshly converted imatrix quants from this PR up as of now if anyone else wants to give them a go: https://huggingface.co/avar6/GLM-5.3-Flash-BF16-gguf For the IQ3_S, I am getting roughly 300t/s prefill at 256K context and 2048 b/ub size. Generation speed starts off at around 9t/s and drops down considerably by mid window (~128K) to around 6t/s. This seems to track with the 'pooled indexer keys' issue. The model is fully coherent and seems to be working fine. I don't have PPL/KL numbers at the moment as I still need to generate a logit dump. I have noticed an interesting memory quirk, which I haven't seen before. This is the only model I have ever seen have inconsistent checkpoint sizes. As the context fills the checkpoints grow alarmingly fast in size. At ~90K they are already up to nearly 1.6GB. I don't know if this is an expected behavior for this model arch, or if this is a something which needs to be looked into. Also, something of note for you @danielhanchen which I found last night while looking over the three PRs for this arch. The vision towers between this PR and yours differ as well. This PR reuses the name clip.vision.projector_type = "glm4v" while you built a new one clip.vision.projector_type = "glm5next". Likely not much of an issue given how easy it is to regenerate mmproj files, but it will need to be delt with as well. |
|
Yes I'll re-do the vision! This is fine! @timkhronos I confirmed timkhronos#9 works fine and does not break your GGUFs. We will however need to do a cheap shard-1 update so that should be fine |
…ed up long context decode, fla, and slight MTP improvements.
|
@timkhronos I saw you changed the tensor naming - but my solution I provided was to allow everyone's quants to work - now your own ones you uploaded don't work haha. We still need to provide the shard rewrite for the naming (glm5-next) which we're fine with, but now the DeepSeek convention means you yourself have to reupload all shards or do a tensor rename inplace with a script - was this your intention? |
|
@danielhanchen Hey! I ended up going with the the indexer_compressor naming scheme, as it is closer to what's already there, and I was meaning to ask Avar to reconvert anyways, as his ggufs were made when we were missing quantization protection for some crucial tensors so they are not ideal. Your vision projectors will need reconverting though most likely, and your main model ggufs might be missing the |
|
@timkhronos Hey! I made some shard rewrites to https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/main/Shard_Rewrite for in preparation! |
| // Whether this context tracks k-pool states. | ||
| bool kpool_track = false; |
There was a problem hiding this comment.
This also seems redundant - no need to duplicate the information when it is directly available from the llama_memory_hybrid_idx * mem. Can be a helper method instead.
There was a problem hiding this comment.
It's not entirely a duplicate, since the flag was also serving as a batch context marker. The update context has mem set too and gets apply()d from kv_self_update, but carries no ubatches to build a state from, so deriving purely from mem would send it into kpool_build_state(get_ubatch()). Setting that flag only in the constructor was an implicit way of ensuring that. Turned it into a helper, with the condition explicit
This comment was marked as spam.
This comment was marked as spam.
|
I tried this PR over RPC. Machine 1 with 4x 9060 XT (GFX1200) and Machine 2 with Strix-Halo (GFX1151) + 7600 XT (GFX1102). All in all ~185GB VRAM to use. I tried the Quant Q3_XL-3.86bpw: At first i could not run it with ROCM or VULKAN. Edit: Just tested VULKAN with -b 256 and -ub 256 there it crashes also. `ROCM -b 1024 -ub 1024: On the ROCM RPC side: Vulkan -b 1024 -ub 1024: RPC side: /home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-rpc/ggml-rpc.cpp:1386: GGML_ASSERT(tensor->data >= buffer_start && tensor->data + tensor_size <= buffer_start + buffer_size) failed` |
|
@timkhronos Heads up, recent change to |
|
@CISC Thank you, adapted in 1b564d2. @ggerganov I ended up adding the multi stream support in ff6be95, so both --kv-unified and --no-kv-unified work properly now. |
…-org#27773 - add LLM_ARCH_GLM5_NEXT (glm5-next) with legacy glm5next alias normalized to new arch - add LLM_KV kpool/select_tail/index_share_mtp and LLM_TENSOR kpool gate/ape - extend hparams with indexer_kpool fields - add LLM_TYPE 320B_A18B and layer kpool tensors - model mapping for both glm5next and glm5-next to new class - model saver, context graph_max_nodes, graph swiglu clamp for GLM5_NEXT - deepseek4 templated mHC helpers for sharing with glm5-next (preserve with_zero_dep) - add src/models/glm5-next.cpp from PR with block_size fallback for legacy GGUF - replace old src/models/glm5next.cpp runtime with alias to new class - kv-cache get_stream/get_k_storage helpers - hybrid idx kpool state merge with checkpoint idx_row_size Source: timkhronos/llama.cpp GLM5.3-Flash 5c4bd50
…ng native ggml-org#27773 path - KDA head dim: primary KDA_HEAD_DIM, fallback SSM_STATE_SIZE - KDA gate lower bound: default -5.0f when missing (legacy semantics) - recurrent layout: prefer explicit RECURRENT_LAYERS, else native n_head_kv==0 if any zero present, else legacy FULL_ATTENTION_INTERVAL default 4
Native ggml-org#27773 with KDA_HEAD_DIM present but no KDA_GATE_LOWER_BOUND should keep -INFINITY (softplus branch). -5.0f applies only when KDA_HEAD_DIM was missing and fell back to SSM_STATE_SIZE (legacy GLM-5.3 format).
|
Data point for the pool-cache discussion (#27773 (comment)) and for anyone sizing this on a single 24 GB card + CPU offload. Setup: EPYC 9355P, 1.5 TB DDR5, 1× RTX 4090, CUDA build. 1. Decode vs depth, this PR vs #27754 (same weights and flags,
Prefill ~123 t/s on both, flat with depth. The gap is the O(n_kv) term the pooled-index cache removes; on a CPU-offloaded box the cache is what makes long context usable. 2. Current head (5c4bd50),
Speeds are the same (the CPU expert stream dominates here). The win of 3. Long context,
No degradation with depth and no sign of the repeat-collapse reported on the other branch; 4. A real agentic session on the same build at |
|
Vision works on 5c4bd50 with UD-Q4_K_XL, Note for anyone hitting "Failed to load CLIP model": the unsloth |
Native ggml-org#27773 metadata (per-layer head_count_kv, kda head dim/gate, indexer kpool/select_tail/index_share_mtp, NoPE MLA, mHC, MoE, swiglu) while keeping the fork direct-quant infrastructure. Adds writer methods, kpool keys, MODEL_TENSORS[GLM5_NEXT], HC/kpool mappings and GLM5V schema keys. Assisted-by: opencode
Add mandatory layer_norm_eps 1e-6, reshape native KDA q/k/v conv1d to the 4D (1, d_inner, 1, d_conv) graph layout in the ordinary path and the same output shape in the direct-quant canonical manifest and direct records via a pure LocalTensor reshape (byte-identical). Sync head_dim cleanup and A_log flatten to n_head. Assisted-by: opencode
…indexer seq_trim Snapshot before replacing the ggml-org#27752-based glm5next port with upstream PR ggml-org#27773. Kept for reference: the last-layer output_norm/output duplicates that keep the trunk graph on the pipeline when the head is pinned to --device-draft, the PIPEDEC_BODY/HEAD graph split, the mtp_only/trunk_only probes for a draft-only GGUF, load_mtp gating for --model-draft, and llama_memory_hybrid_idx::seq_trim. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
Upstream is converging on PR ggml-org#27773 (arch glm5-next) instead; the fork's port of ggml-org#27752 plus the ggml-org#27754 vision graft goes so that PR can be merged as-is. Drops src/models/glm5next.cpp, the converter, the mtmd graft and the glm5next additions to llama-graph / hybrid-idx / model / tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
Upstream's GLM-5.3-Flash port replaces the fork's ggml-org#27752-based one. The one real conflict was two refactors of the same mHC helpers: the fork had moved them into llm_graph_context_dsv4_mla for the SPD/DSpark graphs, the PR into a graph_base<Base> template so glm5-next can stack them on the delta-net base. Resolved by templating the fork's class (llm_graph_context_dsv4_mla_t<Base>, alias llm_graph_context_dsv4_mla) and aliasing the PR's llama_model_deepseek4::graph_base to it, so both the fork's fused-kernel / stream-view helpers and the PR's glm5-next graph compile unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
… trim Re-port of the fork's GLM-5.3-Flash hooks onto the ggml-org#27773 graph: - LLM_GRAPH_TYPE_DECODER_PIPEDEC_BODY ends the trunk at the post-norm hidden state with no out_ids input; graph_pipedec_head is the lm_head over the gathered lane rows on the draft GPU. - With --device-draft pinning the output head, output_norm / output are duplicated onto the last transformer layer so the trunk and body graphs keep ending on the pipeline (the RPC-ends-on-local-backend rule). - A draft-only GGUF (blk.<n_layer>.* + token_embd/output/output_norm, no trunk) loads with the trunk tensors optional, so the 7.4 GiB Q8_0 NextN block can be requantized apart from the trunk and pinned to the GTX 1080. - glm5-next joins the stage-2 / tree eligibility lists. - llama_memory_hybrid_idx::seq_trim marks the pooled-key cache stale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
The ggml-org#27773 port dropped the rollback-plane writes the RS machinery requires: glm5_conv1d wrote only the main conv slot (old graph wrote one snapshot per rollback slot), and the SSM state went through the raw build_delta_net dispatcher plus a direct main-plane write instead of build_recurrent_attn (which emits the K per-token snapshots). After the first partial accept (speculative/DFlash), rollback rounds read never-written planes and the main plane got clobbered with rollback-computed states: drafts degraded (walk top-k -> absurd ids), acceptance collapsed 0.5 -> 0.09 with per-pos decay, output derailed into garbage at ~2-4 t/s. With n_rs_seq == 0 the bug is invisible, which is why target-only stayed clean. Fix: snapshot loop in glm5_conv1d (identical to the slot-0 main write when n_rs_seq == 0) and build_recurrent_attn for the SSM state. Verified: DFlash n_max=7 clean, acceptance 0.46-0.47, 9.2-11.5 t/s; 100-token run stable, matching pre-migration baseline (0.52).
Overview
Add support for GLM -5.3-flash a 320B hybrid model, supporting both text and vision.
Additional information
Architecture
GLM 5.3 flash mixes 34 KDA linear layers with 11 DSA laters, with mHC and Deepseek style Moe. Most of the parts are already in llama.cpp so I reused whatever I could:
What I implemented new:
llama_memory_hybrid_dsa: recurrent state + DSA cache, cloned from hybrid ISWA.Rebased onto llama_memory_hybrid_idx instead of the earlier ISWA clone.the encoder is the same family as glmv4 with per head qk-norm, clamped Swiglu and no post conv norm. It reuses glm4v projector with a swiglu_limit key and an optional image token budget.Added as glm5v as GLM 5.3 Flash requires a different pre processing method than what glm4v uses.Tests
Limitations
Quantized GGUFs converted with this PR are available here.
Requirements