Skip to content

[None][feat] Add GLM-5.3-Flash (glm5_next) support - #19136

Merged
juney-nvidia merged 35 commits into
NVIDIA:mainfrom
ruocheng-nv:feat/glm5.3-flash-full-support
Sep 28, 2026
Merged

juney-nvidia merged 35 commits into
NVIDIA:mainfrom
ruocheng-nv:feat/glm5.3-flash-full-support

Conversation

@ruocheng-nv

@ruocheng-nv ruocheng-nv commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds GLM-5.3-Flash (zai-org/GLM-5.3-Flash, architecture Glm5NextForConditionalGeneration) to the TensorRT-LLM PyTorch backend using the official FP8 checkpoint.

Implements the GLM-5.3-Flash text decoder, native MTP, and vision tower using shared TensorRT-LLM modules. The decoder combines KDA linear attention, sparse MLA with a pool-based indexer, mHC, and MoE.

Shared-module changes:

  • KDA (kimi_kda, KDA decode kernels): low-rank output gates, 64 attention heads, and FP32 gate parameters required by GLM-5.3-Flash, while preserving the existing full-rank gate path. The Kimi K3 KDA MTP decode kernel promotes the state-slot index to Int64 so large state offsets do not overflow (outputs bitwise identical for existing models).
  • mHC: autotuner buckets cover the configured token range; bucket sets are unchanged for existing models at max_num_tokens <= 8192.
  • MLA: the config check allows qk_rope_head_dim == 0, which GLM-5.3-Flash uses; other models are unaffected.
  • Executor: routing to a GLM hybrid cache manager (KDA state + sparse-MLA indexer cache), per-layer head dimensions, multimodal engine and checkpoint-loader hooks.
  • MoE A2A CI unblock: following @zhangcl's recommendation, this PR cherry-picks the NVIDIA kernel-driver gate ported from [None][refactor] overhaul NVLink one-sided all-to-all #19610, so CFT falls back to fence under CUDA forward compatibility. On these systems, a newer user-mode library can expose CFT entry points even when the kernel driver cannot create logical endpoints; without this gate, the reused MoE A2A tests fail in cuLogicalEndpointCreate and abort MPI independently of GLM validation.

Key validated features: TP4/EP4, CUDA graphs, overlap scheduling, chunked prefill, MTP, attention DP (including MTP), FP8 KV cache, prefix reuse with periodic state snapshots, image/video inputs, and text PD disaggregation with and without MTP. Validated on B200.

Documentation: see docs/source/deployment-guide/deployment-guide-for-glm-5.3-flash-on-trtllm.md (serving configs, benchmarking, and the 1K/1K performance chart). Text-only serving uses the Transformers version installed with TensorRT LLM; image and video inputs require transformers==5.17.0. Shared requirements are unchanged.

Release timing: This PR addresses an explicit customer request for GLM-5.3-Flash support.

Test Coverage

Correctness and accuracy

  • Full GSM8K comparisons against HF match within 0.08 percentage points, with and without MTP3, under the matched thinking-mode protocol: 1,319 questions, greedy decoding, and a 4,096-token output budget.
  • Separate full GSM8K regression runs using the repository harness pass for attention DP + MTP3, TP4/EP4 + FP8 KV, and attention DP + MTP3 + FP8 KV. Scores are 93.37%, 93.29%, and 93.86%, respectively. These use the 5-shot raw-completion protocol with a 256-token output budget and are not directly comparable to the HF thinking-mode scores above.
  • Native HF image/video feature parity tests pass with transformers==5.17.0. Full-model validation also covers MMMU, video serving, snapshot-based prefix reuse, chunked prefill, runtime behavior, and NIXL disaggregated serving.
  • Focused unit coverage checks checkpoint routing, layer masks, FP32 gate sharding, encoder-only construction, low-rank KDA projections, sparse-cache layout and dispatch, FP8 cache storage and graph replay, mixed attention-DP/MTP batches, and large KDA state offsets.

CI

Existing unit-test stages are reused (GLM contract tests are added to l0_b200). Full-model coverage adds one TP4/EP4 MTP3 GSM8K case to the existing four-GPU GB300 pre-merge stage (GB300-4_GPUs-PyTorch-1), with CUDA graphs, overlap scheduling, and chunked prefill enabled. Attention DP + MTP3 + FP8 KV and prefix reuse are registered in GB300 multi-GPU post-merge, and MMMU is added to QA. No new Jenkins stages are introduced.

Local resource reference: the MTP3 GSM8K test body took approximately 7.3 minutes on 4x B200, with 8.3 minutes total launcher wall time.

Performance

Measured on 4x B200 with FP8 weights and BF16 KV cache using benchmark_serving: 1,024 input / 1,024 output tokens, random token IDs, seed 0, greedy decoding, ignore_eos, natural MTP acceptance. Each concurrency is warmed up with 2 x C requests, then measured twice with 5 x C requests per run. Output tok/s/user is 1000 / mean TPOT and excludes TTFT; Output tok/s/GPU is aggregate output throughput over the complete request window, including prefill and TTFT, divided by four. The deployment guide chart uses a 2%-tolerant Pareto frontier over the measured TP4 / EP4 and attention DP4 / EP4 configurations (MTP draft length 0, 1, 3): a point is omitted when another point is within 2% of its per-user speed and has higher per-GPU throughput. The table below shows selected frontier points.

C Frontier configuration Output tok/s/user Output tok/s/GPU
1 TP4 / EP4 + MTP3 389 92
4 TP4 / EP4 + MTP3 263 224
16 TP4 / EP4 + MTP3 153 556
64 TP4 / EP4 + MTP1 78 1,150
128 Attention DP4 / EP4 + MTP1 60 1,723
256 Attention DP4 / EP4 + MTP1 46 2,735

Without MTP, TP4 / EP4 reaches 5.57 ms mean TPOT at C=1 (180 tok/s/user) and 6,161 tok/s aggregate output throughput at C=128, so MTP3 improves single-user decode speed by about 2.2x on this synthetic workload. With attention DP, max_batch_size applies per rank (64 per rank for C=256). These measurements cover the tested configurations and concurrency range and do not establish peak throughput or maximum-context performance.

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why.

  • PR follows TRT-LLM Coding Guidelines to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions).

  • If PR introduces API changes, an appropriate PR label is added: api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities.

  • CODEOWNERS is updated if ownership changes.

  • Documentation is updated as needed.

  • Update the architecture diagram if there is a significant design change.

  • The assigned reviewers are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

ruocheng-nv and others added 25 commits September 27, 2026 02:26
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Pass row and head strides into both k-pool scoring branches without copying fused projection views. Cover strided weights and top-512 selection with more than 512 complete pools.

Validation: 12 CPU Triton interpreter cases, offline SM100 compilation, and applicable pre-commit hooks pass. Native GPU execution and full-model accuracy validation remain pending GPU availability.
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Keep the reported performance figures and workload-dependent MTP behavior, while leaving the detailed TTFT investigation in the experiment records. This does not claim the observed TTFT variability is resolved.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Generate and map matching 1024-token tuning buckets above 8192 through the rounded warmup limit, avoiding FMA cache-miss fallbacks for long prefill. Preserve the existing fused-HC policy and cover aligned and non-aligned warmup boundaries with a CPU regression test.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Overlap shared and routed experts under TP CUDA graphs, use native 16-head BF16 sparse MLA for small generation batches, and specialize the FP32 router projection. Preserve existing prefill and FP8 KV attention paths.

Validation: 65 CPU and 27 GPU checks passed, with full GSM8K paired accuracy and MTP3 accuracy/acceptance-length checks. Three-run B200 serving measurements preserve C8 throughput and match local vLLM TP4/EP4 C1 TPOT.
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Widen score row indices before stride multiplication so long-context prefill can address score matrices beyond INT32_MAX elements. Use the shared Torch TopK fallback when CuTe DSL is unavailable, avoiding missing native radix scratch buffers.

Cover all three score-store branches across the 2^31-element boundary and validate candidate bounds and CUDA graph replay without CuTe DSL. Validation: 66 CPU and 37 GPU tests passed, including rebased KDA parity and BF16 state tests.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Remove the GLM-specific small-batch router kernel and its dedicated tests. Use F.linear with FP32 inputs and weights for router logits.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Pass MLA geometry through the shared attention factory, including NoPE support. Move GLM recurrent metadata into its own module, reuse prefer_pinned directly, and guard the module-level FlashMLA import.

Expose raw cache slot indices through KVCacheManagerV2 and replace the GLM private-method call. Cover the shared MLA factory path and extend the existing cache-index test.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Add a keyword-only raw_indices option to the existing V2 cache query, preserving default index conversion for other callers. GLM checks that its sparse layers share one pool and requests raw slots by layer ID.

Remove the separate raw-slot accessor and extend the existing index tests to cover both modes and block-count limits.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Instantiate the optimized KDA schedules for 64 heads so GLM decode remains supported when the new main dispatcher selects a non-compact schedule.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
… access

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Register one optimized MTP3 accuracy case pre-merge and broader post-merge/QA coverage. Add the in-tree config fallback and lazy multimodal processor initialization needed by text tests on the shared Transformers version.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Let Glm5NextCacheManager derive its indexer layers from the attention
layer mask it already receives, so the shared KV cache factory passes
only index_state_dim, and compute the [k | gate | pool key] row width
through one glm_kpool_cache_row_dim helper shared by the factory, the
cache estimator and the model. Rename the factory's is_glm flag to
is_glm5_next, drop its redundant replay-manager check, and restore the
Kimi K3 hybrid cache comments in the shared factory branch.

Drop a greedy stop-word test that duplicates the one already on main,
skip the MMMU test when the native glm5_next processor is unavailable,
point the deployment guide at 1.3.0rc29, scope the transformers 5.17.0
requirement to image and video inputs, and align the GLM file headers
and footnote numbering with the rest of the tree.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Keep only the reason glm5_next rejects non-V2 managers and the Python
NIXL transceiver requirement; the manager's buffer layout is already
documented on Glm5NextCacheManager.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
…lity

Creating a CFT logical endpoint needs kernel-driver support. Under CUDA
forward compatibility the newer user-mode driver still exports the
cuLogicalEndpoint entry points, so the existing gate lets the test run,
and cuLogicalEndpointCreate then fails with CUDA_ERROR_INVALID_VALUE.
The worker aborts MPI_COMM_WORLD and takes down the whole test process.

Skip the CFT cases when the loaded libcuda is newer than the kernel
driver reported by NVML; real CFT failures on a matching driver still
surface.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Forward compatibility pairs a newer user-mode libcuda with an older
kernel driver; compare versions numerically instead of skipping on any
mismatch.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
CFT logical endpoints need kernel-driver support (615+). Under CUDA
forward compatibility a newer user-mode libcuda still exports the
cuLogicalEndpoint entry points, so the existing gates pass, the
workspace is laid out for CFT, and cuLogicalEndpointCreate then fails
with CUDA_ERROR_INVALID_VALUE.

Check the kernel driver version via NVML in resolve_can_use_cft(), which
both workspace sizing and construction use, and fall back to fence when
it is below 615 or cannot be queried. TRTLLM_MOE_A2A_FORCE_CFT=1 does not
bypass this check.

The helpers match the ones in the NVLink one-sided overhaul (NVIDIA#19610) so
that change can take its own copy when it lands. The workspace regression
test now skips its CFT cases through the same helpers instead of parsing
/proc/self/maps.

Signed-off-by: Chulian Zhang <851104+zhangcl@users.noreply.github.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Replace the four-configuration comparison with the Pareto frontier of the
measured TP4/EP4 and attention DP4/EP4 configurations with MTP, drawn in
the style of the GLM-5 guide. Note that attention DP batch size limits
apply per rank, and extend the benchmark concurrency list to match.

Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
@ruocheng-nv
ruocheng-nv force-pushed the feat/glm5.3-flash-full-support branch from 49431d9 to 1ec4cfe Compare September 27, 2026 09:36
@ruocheng-nv

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75486 [ run ] triggered by Bot. Commit: 1ec4cfe Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75486 [ run ] completed with state SUCCESS. Commit: 1ec4cfe
/LLM/main/L0_MergeRequest_PR pipeline #62219 completed with status: 'SUCCESS'
Pipeline passed with automatic retried tests. Check the rerun report for details.

CI Report

Link to invocation

@juney-nvidia
juney-nvidia merged commit d21a8ce into NVIDIA:main Sep 28, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.