[None][feat] Add GLM-5.3-Flash (glm5_next) support - #19136
Merged
juney-nvidia merged 35 commits intoSep 28, 2026
Merged
Conversation
ruocheng-nv
requested review from
2ez4bz,
Mgluhovskoi,
WeiHaocheng,
chienchunhung,
nv-xtf,
pengbowang-nv,
rosong11,
thorjohnsen,
yizhang-nv and
yunruis
September 14, 2026 05:54
ruocheng-nv
requested review from
BowenFu,
asfiyab-nvidia,
brnguyen2,
crazydemo,
hanjingtian,
weiminwang-nv,
zhaoyuanh-nvidia and
zhenhuaw-me
September 14, 2026 05:54
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Pass row and head strides into both k-pool scoring branches without copying fused projection views. Cover strided weights and top-512 selection with more than 512 complete pools. Validation: 12 CPU Triton interpreter cases, offline SM100 compilation, and applicable pre-commit hooks pass. Native GPU execution and full-model accuracy validation remain pending GPU availability. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Keep the reported performance figures and workload-dependent MTP behavior, while leaving the detailed TTFT investigation in the experiment records. This does not claim the observed TTFT variability is resolved. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Generate and map matching 1024-token tuning buckets above 8192 through the rounded warmup limit, avoiding FMA cache-miss fallbacks for long prefill. Preserve the existing fused-HC policy and cover aligned and non-aligned warmup boundaries with a CPU regression test. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Overlap shared and routed experts under TP CUDA graphs, use native 16-head BF16 sparse MLA for small generation batches, and specialize the FP32 router projection. Preserve existing prefill and FP8 KV attention paths. Validation: 65 CPU and 27 GPU checks passed, with full GSM8K paired accuracy and MTP3 accuracy/acceptance-length checks. Three-run B200 serving measurements preserve C8 throughput and match local vLLM TP4/EP4 C1 TPOT. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Widen score row indices before stride multiplication so long-context prefill can address score matrices beyond INT32_MAX elements. Use the shared Torch TopK fallback when CuTe DSL is unavailable, avoiding missing native radix scratch buffers. Cover all three score-store branches across the 2^31-element boundary and validate candidate bounds and CUDA graph replay without CuTe DSL. Validation: 66 CPU and 37 GPU tests passed, including rebased KDA parity and BF16 state tests. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Remove the GLM-specific small-batch router kernel and its dedicated tests. Use F.linear with FP32 inputs and weights for router logits. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Pass MLA geometry through the shared attention factory, including NoPE support. Move GLM recurrent metadata into its own module, reuse prefer_pinned directly, and guard the module-level FlashMLA import. Expose raw cache slot indices through KVCacheManagerV2 and replace the GLM private-method call. Cover the shared MLA factory path and extend the existing cache-index test. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Add a keyword-only raw_indices option to the existing V2 cache query, preserving default index conversion for other callers. GLM checks that its sparse layers share one pool and requests raw slots by layer ID. Remove the separate raw-slot accessor and extend the existing index tests to cover both modes and block-count limits. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Instantiate the optimized KDA schedules for 64 heads so GLM decode remains supported when the new main dispatcher selects a non-compact schedule. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
… access Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Register one optimized MTP3 accuracy case pre-merge and broader post-merge/QA coverage. Add the in-tree config fallback and lazy multimodal processor initialization needed by text tests on the shared Transformers version. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Let Glm5NextCacheManager derive its indexer layers from the attention layer mask it already receives, so the shared KV cache factory passes only index_state_dim, and compute the [k | gate | pool key] row width through one glm_kpool_cache_row_dim helper shared by the factory, the cache estimator and the model. Rename the factory's is_glm flag to is_glm5_next, drop its redundant replay-manager check, and restore the Kimi K3 hybrid cache comments in the shared factory branch. Drop a greedy stop-word test that duplicates the one already on main, skip the MMMU test when the native glm5_next processor is unavailable, point the deployment guide at 1.3.0rc29, scope the transformers 5.17.0 requirement to image and video inputs, and align the GLM file headers and footnote numbering with the rest of the tree. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Keep only the reason glm5_next rejects non-V2 managers and the Python NIXL transceiver requirement; the manager's buffer layout is already documented on Glm5NextCacheManager. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
…lity Creating a CFT logical endpoint needs kernel-driver support. Under CUDA forward compatibility the newer user-mode driver still exports the cuLogicalEndpoint entry points, so the existing gate lets the test run, and cuLogicalEndpointCreate then fails with CUDA_ERROR_INVALID_VALUE. The worker aborts MPI_COMM_WORLD and takes down the whole test process. Skip the CFT cases when the loaded libcuda is newer than the kernel driver reported by NVML; real CFT failures on a matching driver still surface. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Forward compatibility pairs a newer user-mode libcuda with an older kernel driver; compare versions numerically instead of skipping on any mismatch. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
CFT logical endpoints need kernel-driver support (615+). Under CUDA forward compatibility a newer user-mode libcuda still exports the cuLogicalEndpoint entry points, so the existing gates pass, the workspace is laid out for CFT, and cuLogicalEndpointCreate then fails with CUDA_ERROR_INVALID_VALUE. Check the kernel driver version via NVML in resolve_can_use_cft(), which both workspace sizing and construction use, and fall back to fence when it is below 615 or cannot be queried. TRTLLM_MOE_A2A_FORCE_CFT=1 does not bypass this check. The helpers match the ones in the NVLink one-sided overhaul (NVIDIA#19610) so that change can take its own copy when it lands. The workspace regression test now skips its CFT cases through the same helpers instead of parsing /proc/self/maps. Signed-off-by: Chulian Zhang <851104+zhangcl@users.noreply.github.com> Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
Replace the four-configuration comparison with the Pareto frontier of the measured TP4/EP4 and attention DP4/EP4 configurations with MTP, drawn in the style of the GLM-5 guide. Note that attention DP batch size limits apply per rank, and extend the benchmark concurrency list to match. Signed-off-by: Ruocheng Jia <ruochengj@nvidia.com>
ruocheng-nv
force-pushed
the
feat/glm5.3-flash-full-support
branch
from
September 27, 2026 09:36
49431d9 to
1ec4cfe
Compare
Contributor
Author
|
/bot run --disable-fail-fast |
Collaborator
|
PR_Github #75486 [ run ] triggered by Bot. Commit: |
Collaborator
|
PR_Github #75486 [ run ] completed with state |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds GLM-5.3-Flash (
zai-org/GLM-5.3-Flash, architectureGlm5NextForConditionalGeneration) to the TensorRT-LLM PyTorch backend using the official FP8 checkpoint.Implements the GLM-5.3-Flash text decoder, native MTP, and vision tower using shared TensorRT-LLM modules. The decoder combines KDA linear attention, sparse MLA with a pool-based indexer, mHC, and MoE.
Shared-module changes:
kimi_kda, KDA decode kernels): low-rank output gates, 64 attention heads, and FP32 gate parameters required by GLM-5.3-Flash, while preserving the existing full-rank gate path. The Kimi K3 KDA MTP decode kernel promotes the state-slot index to Int64 so large state offsets do not overflow (outputs bitwise identical for existing models).max_num_tokens<= 8192.qk_rope_head_dim == 0, which GLM-5.3-Flash uses; other models are unaffected.cuLogicalEndpointCreateand abort MPI independently of GLM validation.Key validated features: TP4/EP4, CUDA graphs, overlap scheduling, chunked prefill, MTP, attention DP (including MTP), FP8 KV cache, prefix reuse with periodic state snapshots, image/video inputs, and text PD disaggregation with and without MTP. Validated on B200.
Documentation: see
docs/source/deployment-guide/deployment-guide-for-glm-5.3-flash-on-trtllm.md(serving configs, benchmarking, and the 1K/1K performance chart). Text-only serving uses the Transformers version installed with TensorRT LLM; image and video inputs requiretransformers==5.17.0. Shared requirements are unchanged.Release timing: This PR addresses an explicit customer request for GLM-5.3-Flash support.
Test Coverage
Correctness and accuracy
transformers==5.17.0. Full-model validation also covers MMMU, video serving, snapshot-based prefix reuse, chunked prefill, runtime behavior, and NIXL disaggregated serving.CI
Existing unit-test stages are reused (GLM contract tests are added to
l0_b200). Full-model coverage adds one TP4/EP4 MTP3 GSM8K case to the existing four-GPU GB300 pre-merge stage (GB300-4_GPUs-PyTorch-1), with CUDA graphs, overlap scheduling, and chunked prefill enabled. Attention DP + MTP3 + FP8 KV and prefix reuse are registered in GB300 multi-GPU post-merge, and MMMU is added to QA. No new Jenkins stages are introduced.Local resource reference: the MTP3 GSM8K test body took approximately 7.3 minutes on 4x B200, with 8.3 minutes total launcher wall time.
Performance
Measured on 4x B200 with FP8 weights and BF16 KV cache using
benchmark_serving: 1,024 input / 1,024 output tokens, random token IDs, seed 0, greedy decoding,ignore_eos, natural MTP acceptance. Each concurrency is warmed up with 2 x C requests, then measured twice with 5 x C requests per run.Output tok/s/useris1000 / mean TPOTand excludes TTFT;Output tok/s/GPUis aggregate output throughput over the complete request window, including prefill and TTFT, divided by four. The deployment guide chart uses a 2%-tolerant Pareto frontier over the measured TP4 / EP4 and attention DP4 / EP4 configurations (MTP draft length 0, 1, 3): a point is omitted when another point is within 2% of its per-user speed and has higher per-GPU throughput. The table below shows selected frontier points.Without MTP, TP4 / EP4 reaches 5.57 ms mean TPOT at C=1 (180 tok/s/user) and 6,161 tok/s aggregate output throughput at C=128, so MTP3 improves single-user decode speed by about 2.2x on this synthetic workload. With attention DP,
max_batch_sizeapplies per rank (64 per rank for C=256). These measurements cover the tested configurations and concurrency range and do not establish peak throughput or maximum-context performance.PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why.
PR follows TRT-LLM Coding Guidelines to the best of your knowledge.
Test cases are provided for new code paths (see test instructions).
If PR introduces API changes, an appropriate PR label is added:
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities.
CODEOWNERS is updated if ownership changes.
Documentation is updated as needed.
Update the architecture diagram if there is a significant design change.
The assigned reviewers are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.