Skip to content

add GLM-5.3-Flash (GLM5-Next) support - #27773

Open
timkhronos wants to merge 41 commits into
ggml-org:masterfrom
timkhronos:GLM5.3-Flash
Open

add GLM-5.3-Flash (GLM5-Next) support#27773
timkhronos wants to merge 41 commits into
ggml-org:masterfrom
timkhronos:GLM5.3-Flash

Conversation

@timkhronos

@timkhronos timkhronos commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Overview

Add support for GLM -5.3-flash a 320B hybrid model, supporting both text and vision.

Additional information

Architecture

GLM 5.3 flash mixes 34 KDA linear layers with 11 DSA laters, with mHC and Deepseek style Moe. Most of the parts are already in llama.cpp so I reused whatever I could:

  • KDA layers reuse the Kimi-K3 implementation
  • Attention layers are nope only MLA
  • For mHC I reused the Deepseek V4 implementation. I moved the build_hc helpers from the DSV4 graph into graph_context so both models can share them.
  • Moe and swiglu clamping follow DSV4.

What I implemented new:

  • Here, the DSA indexer scores pools of 4 consecutive token, and always keeps the incomplete tail. I implemented this on top of the existing DSA cache. The indexer cache stores key | gate per token and pooling happens in the graph. No new backend ops have been added.
  • llama_memory_hybrid_dsa: recurrent state + DSA cache, cloned from hybrid ISWA. Rebased onto llama_memory_hybrid_idx instead of the earlier ISWA clone.
  • Vision Tower: the encoder is the same family as glmv4 with per head qk-norm, clamped Swiglu and no post conv norm. It reuses glm4v projector with a swiglu_limit key and an optional image token budget. Added as glm5v as GLM 5.3 Flash requires a different pre processing method than what glm4v uses.
  • Small, precision sensitive tensors (indexer, mHC mixers, KDA gates, MLA low rank paths, roughly 1GB total) are kept unquantized.

Tests

  • Logits match transformers on a small random model across full prefill, small ubatches and single token decode while sparse selection is active, covering both scatter and gather.
  • Vision embeddings match to ~1e-5.
  • The converted model generates coherently and correctly, and vision is working as expected.

Limitations

Quantized GGUFs converted with this PR are available here.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, AI was used in an assistive capacity, and helped figure out and solve several conversion issues, and helped validate the correctness of the implementation.

@github-actions github-actions Bot added model Model specific mtmd Related to multimodal functionality (video/image/audio) conversion labels Aug 26, 2026
Comment thread tools/mtmd/clip.cpp Outdated
Comment thread tools/mtmd/clip-model.h Outdated
Comment thread tools/mtmd/clip-impl.h Outdated
@danielhanchen

danielhanchen commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

Hey @timkhronos great work on the PR! A few requests if possible:

  1. You're using glm5-next for general.architecture, whilst model: add GLM-5-Next (GLM-5.3-Flash) #27754 and model : add GLM-5.3-Flash (glm5next) #27752 uses glm5next - Qwen3-Next for eg does qwen3next - I'm unsure what the convention is @ngxson but adopting glm5next might be more generalized? If glm5-next is accepted, a simple first shard rewrite for our uploads should suffice.

  2. The bigger issue is blk.N.indexer.kpool_ape / kpool_gate vs blk.N.indexer_compressor_ape / _gate. deepseek4 mainline already uses blk.N.indexer_compressor_ape and blk.N.indexer_compressor_gate but your PR changes it - the ones we uploaded uses deepseek4's convention. If this PR is accepted, we have to provide a script to rewrite all tensor names or folks have to re-download. If not, can you add aliases so the ones we published works - thanks in advance. Seems like it's more complex than I expected.

Tagging @ngxson for visibility as well.

I re-checked and if (1) + (2) is applied, the quants we uploaded work fine (+ the small shard-1 rewrite) and KLD / PPL are correct under this PR.

Seems like a simple alias isn't possible actually :( It breaks the quants made with this PR

@danielhanchen

Copy link
Copy Markdown
Contributor

Hmmm https://github.com/timkhronos/llama.cpp/pull/9/changes would alias the tensors but it looks a bit problematic hmmm

@Sciguy429

Copy link
Copy Markdown

Throwing up some performance numbers here from the lower end of consumer hardware (128GB DDR5 + 24GB VRAM (4090)).

avar6 has some freshly converted imatrix quants from this PR up as of now if anyone else wants to give them a go: https://huggingface.co/avar6/GLM-5.3-Flash-BF16-gguf

For the IQ3_S, I am getting roughly 300t/s prefill at 256K context and 2048 b/ub size. Generation speed starts off at around 9t/s and drops down considerably by mid window (~128K) to around 6t/s. This seems to track with the 'pooled indexer keys' issue. The model is fully coherent and seems to be working fine. I don't have PPL/KL numbers at the moment as I still need to generate a logit dump.

I have noticed an interesting memory quirk, which I haven't seen before. This is the only model I have ever seen have inconsistent checkpoint sizes. As the context fills the checkpoints grow alarmingly fast in size. At ~90K they are already up to nearly 1.6GB. I don't know if this is an expected behavior for this model arch, or if this is a something which needs to be looked into.

Also, something of note for you @danielhanchen which I found last night while looking over the three PRs for this arch. The vision towers between this PR and yours differ as well. This PR reuses the name clip.vision.projector_type = "glm4v" while you built a new one clip.vision.projector_type = "glm5next". Likely not much of an issue given how easy it is to regenerate mmproj files, but it will need to be delt with as well.

@danielhanchen

Copy link
Copy Markdown
Contributor

Yes I'll re-do the vision! This is fine!

@timkhronos I confirmed timkhronos#9 works fine and does not break your GGUFs. We will however need to do a cheap shard-1 update so that should be fine

@danielhanchen

Copy link
Copy Markdown
Contributor

@timkhronos I saw you changed the tensor naming - but my solution I provided was to allow everyone's quants to work - now your own ones you uploaded don't work haha.

We still need to provide the shard rewrite for the naming (glm5-next) which we're fine with, but now the DeepSeek convention means you yourself have to reupload all shards or do a tensor rename inplace with a script - was this your intention?

@timkhronos

timkhronos commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

@danielhanchen Hey!

I ended up going with the the indexer_compressor naming scheme, as it is closer to what's already there, and I was meaning to ask Avar to reconvert anyways, as his ggufs were made when we were missing quantization protection for some crucial tensors so they are not ideal.

Your vision projectors will need reconverting though most likely, and your main model ggufs might be missing the index_share_for_mtp_iteration key as well.

@danielhanchen

Copy link
Copy Markdown
Contributor

@timkhronos Hey! I made some shard rewrites to https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/main/Shard_Rewrite for in preparation!

Comment thread src/llama-memory-hybrid-idx.h Outdated
Comment on lines +185 to +186
// Whether this context tracks k-pool states.
bool kpool_track = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This also seems redundant - no need to duplicate the information when it is directly available from the llama_memory_hybrid_idx * mem. Can be a helper method instead.

@timkhronos timkhronos Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not entirely a duplicate, since the flag was also serving as a batch context marker. The update context has mem set too and gets apply()d from kv_self_update, but carries no ubatches to build a state from, so deriving purely from mem would send it into kpool_build_state(get_ubatch()). Setting that flag only in the constructor was an implicit way of ensuring that. Turned it into a helper, with the condition explicit

Comment thread src/llama-context.cpp Outdated
@miminashi

This comment was marked as spam.

@Neresco

Neresco commented Sep 2, 2026

Copy link
Copy Markdown

I tried this PR over RPC. Machine 1 with 4x 9060 XT (GFX1200) and Machine 2 with Strix-Halo (GFX1151) + 7600 XT (GFX1102). All in all ~185GB VRAM to use.

I tried the Quant Q3_XL-3.86bpw:
https://huggingface.co/avar6/GLM-5.3-Flash-BF16-gguf/tree/main/Q3_XL-3.86bpw

At first i could not run it with ROCM or VULKAN.
I had success as i lowered the -b and -ub both to 128 and retried.
This works for VULKAN but ROCM still chrashes even with -b and -ub 128.

Edit: Just tested VULKAN with -b 256 and -ub 256 there it crashes also.
Same Terminal crash output like below.

`ROCM -b 1024 -ub 1024:
13.28.075.747 I slot get_availabl: id 3 | task -1 | selected slot by LRU, t_last = -1
13.28.075.792 I slot launch_slot_: id 3 | task 0 | processing task, is_child = 0
13.44.644.302 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2048, progress = 0.58, t = 16.57 s / 123.63 tokens per second
13.52.383.137 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 2988, progress = 0.85, t = 24.30 s / 122.94 tokens per second
13.57.068.995 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 3500, progress = 1.00, t = 28.99 s / 120.73 tokens per second
14.02.350.461 E recv failed (bytes_recv=0, size_to_recv=8)
/home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-rpc/ggml-rpc.cpp:566: Remote RPC server crashed or returned malformed response

On the ROCM RPC side:
ggml_cuda_compute_forward: MUL_MAT failed
ROCm error: no kernel image is available for execution on the device
current device: 1, in function ggml_cuda_compute_forward at /home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-cuda/ggml-cuda.cu:2416
err
/home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-cuda/ggml-cuda.cu:107: ROCm error

Vulkan -b 1024 -ub 1024:
7.11.282.030 I slot print_timing: id 3 | task 0 | prompt processing, n_tokens = 1024, progress = 0.29, t = 29.83 s / 34.33 tokens per second
/home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-rpc/ggml-rpc.cpp:566: Remote RPC server crashed or returned malformed response
7.14.483.216 E recv failed (bytes_recv=0, size_to_recv=8)

RPC side:

/home/lunarbuntu/Downloads/llama.cpp-GLM5.3-Flash/ggml/src/ggml-rpc/ggml-rpc.cpp:1386: GGML_ASSERT(tensor->data >= buffer_start && tensor->data + tensor_size <= buffer_start + buffer_size) failed`

@CISC

CISC commented Sep 3, 2026

Copy link
Copy Markdown
Member

@timkhronos Heads up, recent change to n_ff_exp requires rebase and adaptation.

@timkhronos

Copy link
Copy Markdown
Contributor Author

@CISC Thank you, adapted in 1b564d2.

@ggerganov I ended up adding the multi stream support in ff6be95, so both --kv-unified and --no-kv-unified work properly now.

NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 4, 2026
…-org#27773

- add LLM_ARCH_GLM5_NEXT (glm5-next) with legacy glm5next alias normalized to new arch
- add LLM_KV kpool/select_tail/index_share_mtp and LLM_TENSOR kpool gate/ape
- extend hparams with indexer_kpool fields
- add LLM_TYPE 320B_A18B and layer kpool tensors
- model mapping for both glm5next and glm5-next to new class
- model saver, context graph_max_nodes, graph swiglu clamp for GLM5_NEXT
- deepseek4 templated mHC helpers for sharing with glm5-next (preserve with_zero_dep)
- add src/models/glm5-next.cpp from PR with block_size fallback for legacy GGUF
- replace old src/models/glm5next.cpp runtime with alias to new class
- kv-cache get_stream/get_k_storage helpers
- hybrid idx kpool state merge with checkpoint idx_row_size

Source: timkhronos/llama.cpp GLM5.3-Flash 5c4bd50
NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 4, 2026
…ng native ggml-org#27773 path

- KDA head dim: primary KDA_HEAD_DIM, fallback SSM_STATE_SIZE
- KDA gate lower bound: default -5.0f when missing (legacy semantics)
- recurrent layout: prefer explicit RECURRENT_LAYERS, else native n_head_kv==0 if any zero present, else legacy FULL_ATTENTION_INTERVAL default 4
NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 4, 2026
Native ggml-org#27773 with KDA_HEAD_DIM present but no KDA_GATE_LOWER_BOUND
should keep -INFINITY (softplus branch). -5.0f applies only when
KDA_HEAD_DIM was missing and fell back to SSM_STATE_SIZE (legacy
GLM-5.3 format).
@nicholasshirley

nicholasshirley commented Sep 4, 2026

Copy link
Copy Markdown

Data point for the pool-cache discussion (#27773 (comment)) and for anyone sizing this on a single 24 GB card + CPU offload.

Setup: EPYC 9355P, 1.5 TB DDR5, 1× RTX 4090, CUDA build. -ngl 99 with the routed+shared experts, router and token_embd overridden to CPU; attention, indexer, MLA low-rank paths and output on the GPU. unsloth/GLM-5.3-Flash-GGUF UD-Q4_K_XL with the Shard_Rewrite/ shard-1, greedy, -b 4096 --parallel 1, NVIDIA_TF32_OVERRIDE=0.

1. Decode vs depth, this PR vs #27754 (same weights and flags, -fa off -ub 1024 -c 32768, 2026-08-31 builds: 27773@a771613, 27754 head of that day)

depth #27754 tg t/s #27773 tg t/s
~20 24.8 25.3
7.5k 21.1 24.0
28.8k 14.6 23.7

Prefill ~123 t/s on both, flat with depth. The gap is the O(n_kv) term the pooled-index cache removes; on a CPU-offloaded box the cache is what makes long context usable.

2. Current head (5c4bd50), -fa off vs -fa on (-ub 1024 -c 32768, greedy)

depth -fa off pp / tg t/s VRAM peak -fa on pp / tg t/s VRAM peak
~10 48 / 25.0 17.4 GB 49 / 25.4 10.9 GB
7.5k 122 / 23.9 18.4 GB 124 / 24.0 10.9 GB
28.8k 123 / 23.8 21.1 GB 130 / 23.8 10.9 GB

Speeds are the same (the CPU expert stream dominates here). The win of -fa on is memory: the fa-off path pads V to 512 and its compute buffer grows with n_kv, while fa-on stays flat at 10.9 GB, which is what lets >100k context fit on 24 GB at -ub 1024. -fa off cannot allocate its compute buffers at -c 131072 even at -ub 512.

3. Long context, -fa on -ub 1024 -c 131072 (12.6 GB at load). Needle retrieval in chat format: wikitext with two planted "secret code word" sentences at 2 % and 98 % of the text, then two questions.

prompt tokens pp t/s tg t/s VRAM peak both needles last-paragraph question
29.9k 132 23.7 12.8 GB yes correct
45.5k 137 23.6 12.8 GB yes correct
58.0k 138 23.5 12.8 GB yes correct
92.5k 137 23.4 12.8 GB yes correct

No degradation with depth and no sign of the repeat-collapse reported on the other branch; -ub 128 and -ub 1024 gave byte-identical greedy output at 58k. One caveat for anyone probing this: a raw /completion greedy continuation of wikitext loops into repetition from ~29k tokens on, identically at both ubatch sizes, so that is the sampling regime, not the graph. The chat-format retrieval test is the one that measures the model.

4. A real agentic session on the same build at -c 524288 (20.1 GB at load): an opencode coding agent reviewing an 81-file PR, 30 requests, 42.8k generated tokens, context growing to 160k by appending tool output. Decode 23.3 t/s at 13k → 22.3 t/s at 150k, prefill 120-138 t/s throughout, no coherence problems at any depth.

@nicholasshirley

Copy link
Copy Markdown

Vision works on 5c4bd50 with UD-Q4_K_XL, -fa on, ctx 524288, projector on CPU. 1500x500 logo -> 1003 prompt tokens, 2875x1500 diagram -> 5586 tokens, labels read correctly.

Note for anyone hitting "Failed to load CLIP model": the unsloth Shard_Rewrite mmproj predates the swiglu_clamp rename; avar6's mmproj has the current keys.

NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 4, 2026
Native ggml-org#27773 metadata (per-layer head_count_kv, kda head dim/gate,
indexer kpool/select_tail/index_share_mtp, NoPE MLA, mHC, MoE, swiglu)
while keeping the fork direct-quant infrastructure. Adds writer methods,
kpool keys, MODEL_TENSORS[GLM5_NEXT], HC/kpool mappings and GLM5V schema keys.

Assisted-by: opencode
NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 4, 2026
Add mandatory layer_norm_eps 1e-6, reshape native KDA q/k/v conv1d to the
4D (1, d_inner, 1, d_conv) graph layout in the ordinary path and the same
output shape in the direct-quant canonical manifest and direct records via a
pure LocalTensor reshape (byte-identical). Sync head_dim cleanup and A_log
flatten to n_head.

Assisted-by: opencode
MarkShark2 added a commit to MarkShark2/llama.cpp that referenced this pull request Sep 5, 2026
…indexer seq_trim

Snapshot before replacing the ggml-org#27752-based glm5next port with upstream PR ggml-org#27773.
Kept for reference: the last-layer output_norm/output duplicates that keep the
trunk graph on the pipeline when the head is pinned to --device-draft, the
PIPEDEC_BODY/HEAD graph split, the mtp_only/trunk_only probes for a draft-only
GGUF, load_mtp gating for --model-draft, and llama_memory_hybrid_idx::seq_trim.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
MarkShark2 added a commit to MarkShark2/llama.cpp that referenced this pull request Sep 5, 2026
Upstream is converging on PR ggml-org#27773 (arch glm5-next) instead; the fork's
port of ggml-org#27752 plus the ggml-org#27754 vision graft goes so that PR can be merged
as-is. Drops src/models/glm5next.cpp, the converter, the mtmd graft and
the glm5next additions to llama-graph / hybrid-idx / model / tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
MarkShark2 added a commit to MarkShark2/llama.cpp that referenced this pull request Sep 5, 2026
Upstream's GLM-5.3-Flash port replaces the fork's ggml-org#27752-based one. The one
real conflict was two refactors of the same mHC helpers: the fork had moved
them into llm_graph_context_dsv4_mla for the SPD/DSpark graphs, the PR into a
graph_base<Base> template so glm5-next can stack them on the delta-net base.
Resolved by templating the fork's class (llm_graph_context_dsv4_mla_t<Base>,
alias llm_graph_context_dsv4_mla) and aliasing the PR's
llama_model_deepseek4::graph_base to it, so both the fork's fused-kernel /
stream-view helpers and the PR's glm5-next graph compile unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
MarkShark2 added a commit to MarkShark2/llama.cpp that referenced this pull request Sep 5, 2026
… trim

Re-port of the fork's GLM-5.3-Flash hooks onto the ggml-org#27773 graph:
- LLM_GRAPH_TYPE_DECODER_PIPEDEC_BODY ends the trunk at the post-norm hidden
  state with no out_ids input; graph_pipedec_head is the lm_head over the
  gathered lane rows on the draft GPU.
- With --device-draft pinning the output head, output_norm / output are
  duplicated onto the last transformer layer so the trunk and body graphs
  keep ending on the pipeline (the RPC-ends-on-local-backend rule).
- A draft-only GGUF (blk.<n_layer>.* + token_embd/output/output_norm, no
  trunk) loads with the trunk tensors optional, so the 7.4 GiB Q8_0 NextN
  block can be requantized apart from the trunk and pinned to the GTX 1080.
- glm5-next joins the stage-2 / tree eligibility lists.
- llama_memory_hybrid_idx::seq_trim marks the pooled-key cache stale.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
NovNovikov added a commit to NovNovikov/llama.cpp that referenced this pull request Sep 5, 2026
The ggml-org#27773 port dropped the rollback-plane writes the RS machinery
requires: glm5_conv1d wrote only the main conv slot (old graph wrote
one snapshot per rollback slot), and the SSM state went through the
raw build_delta_net dispatcher plus a direct main-plane write instead
of build_recurrent_attn (which emits the K per-token snapshots).

After the first partial accept (speculative/DFlash), rollback rounds
read never-written planes and the main plane got clobbered with
rollback-computed states: drafts degraded (walk top-k -> absurd ids),
acceptance collapsed 0.5 -> 0.09 with per-pos decay, output derailed
into garbage at ~2-4 t/s. With n_rs_seq == 0 the bug is invisible,
which is why target-only stayed clean.

Fix: snapshot loop in glm5_conv1d (identical to the slot-0 main write
when n_rs_seq == 0) and build_recurrent_attn for the SSM state.
Verified: DFlash n_max=7 clean, acceptance 0.46-0.47, 9.2-11.5 t/s;
100-token run stable, matching pre-migration baseline (0.52).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion model Model specific mtmd Related to multimodal functionality (video/image/audio) testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants