Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
372 commits
Select commit Hold shift + click to select a range
c22a2c0
[None][test] update coderabbit prompt (#16478)
xinhe-nv Jul 21, 2026
58298ad
[TRTLLM-12838][infra] CBTS coverage db audit (#16658)
crazydemo Jul 21, 2026
306c8ce
[None][infra] Waive 2 failed cases for main in pre-merge 30273 (#16675)
trtllm-agent Jul 21, 2026
257c8e0
[TRTLLM-13117][feat] Implement Uneven TP Linear for VisualGen models …
belgarten-nv Jul 21, 2026
a75e333
[None][infra] Add per-stage opt to cap/disable stage-level infra retr…
dpitman-nvda Jul 21, 2026
c988141
[None][infra] Waive 2 failed cases for main in pre-merge 48918 (#16682)
trtllm-agent Jul 21, 2026
6fabfdc
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jul 21, 2026
412768d
[None][infra] Waive 21 failed cases for main in post-merge 2850 (#16700)
trtllm-agent Jul 22, 2026
56f28cd
[TRTLLM-14027][infra] Remove --trt_root and stop installing the Tenso…
Wanli-Jiang Jul 22, 2026
8e2816b
[None][infra] Waive 1 failed cases for main in pre-merge 49101 (#16698)
trtllm-agent Jul 22, 2026
09e5d0c
[None][perf] Size the per-tensor FP8 dynamic-quant amax grid to the i…
hyukn Jul 22, 2026
ea504d1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jul 22, 2026
9714902
[None][feat] Raise the CuTE-DSL top-k decode limit to 16384 and suppo…
Hudayday Jul 22, 2026
4610a2a
[None][perf] Add batched-pybind fast-path in TorchSampler.update_requ…
chenfeiz0326 Jul 22, 2026
09e9738
[TRTLLM-14510][feat] serve: multi-process HTTP frontends on the class…
lancelly Jul 22, 2026
697738c
[https://nvbugs/6438658][fix] Fix KV cache estimation capacity (#16545)
jiaganc Jul 22, 2026
128d020
[None][fix] Make auto host tier sizing rank-aware in KVCacheManagerV2…
erictsai-nv Jul 22, 2026
858fd17
[None][feat] Support Nemotron dynamic-tree MTP decoding (#15582)
sunnyqgg Jul 22, 2026
c54538b
[https://nvbugs/6480574][fix] seed attrs bypassed by __new__ in pool …
JunyiXu-nv Jul 22, 2026
9095cc1
[None][perf] DSv4 GVR top-k (CuTe DSL): P4 histogram refinement + red…
siyidNV Jul 22, 2026
7965292
[https://nvbugs/6468821][infra] Split slow B300 attention unit tests …
yuxianq Jul 22, 2026
9cd6a49
[https://nvbugs/6422318][fix] Cast cos_sin_cache to float32 at cat-ti…
trtllm-agent Jul 22, 2026
f9c253e
[TRTLLM-14473][chore] Remove legacy TensorRT backend tests examples a…
Wanli-Jiang Jul 22, 2026
45d1c11
[TRTLLM-14019][feat] Add MiniMax-M3 MSA sparse attention backend [rev…
brb-nv Jul 22, 2026
8341fb1
[#12595][feat] Emit initial KV cache stats at startup for external me…
BenjaminBraunDev Jul 22, 2026
9e96f8b
[TRTLLM-14255][fix] migrate MiniMax M3 to loader v2 for TP8 support (…
peihu-nv Jul 22, 2026
5e87763
[TRTLLM-13231][feat] Support top_p_decay in the PyTorch TorchSampler …
zhaoyangwang-nvidia Jul 22, 2026
1fbd240
[https://nvbugs/6442074][fix] Rename MTPEagleDynamicTreeWorker.forwar…
sunnyqgg Jul 22, 2026
b294868
[None][feat] Add BaseMultimodalDummyInputsBuilder to minimaxm3_vl (#1…
pcicotti Jul 22, 2026
c1483b9
[TRTLLMINF-102][fix] Surface SLURM device faults to the failure class…
dpitman-nvda Jul 22, 2026
1d2e79e
[https://nvbugs/6499882][fix] Seed _multi_frontend_ipc_dir in bare pr…
pranav-nvidia Jul 23, 2026
7418b69
[None][infra] Preview/bump/main (#16758)
yuanjingx87 Jul 23, 2026
69c3cf8
[TRTLLM-13694][feat] Add IBDB recipe provenance and refresh configs (…
Mgluhovskoi Jul 23, 2026
a19410e
[TRTLLM-14540][perf] Skip fp32 state round-trip in FlashInfer GDN pre…
nv-guomingz Jul 23, 2026
0a401b4
[None][chore] Remove attention backend test waivers (#16723)
yuxianq Jul 23, 2026
b039094
[#11932][fix] Filter CUTLASS MoE GEMM tile configs by device shared m…
mihai-chiorean Jul 23, 2026
8041493
[TRTLLM-13969][feat] Support MiniMax M3 for Disaggregated Serving (#1…
peihu-nv Jul 23, 2026
3d86c72
[TRTLLMINF-188][infra] Require approval for PerfSanity wildcard runs …
chzblych Jul 23, 2026
c6e0c98
[None][infra] Waive 1 failed cases for main in pre-merge 49424 (#16781)
trtllm-agent Jul 23, 2026
3867b2a
[None][infra] Waive 1 failed cases for main in pre-merge 49424 (#16780)
trtllm-agent Jul 23, 2026
93efd33
[None][feat] ADP conversation router: configurable least-queued place…
lancelly Jul 23, 2026
f52f3fe
[None][fix] Load DeepSeek V4 mixed-precision NVFP4 checkpoints (#16433)
lfr-0531 Jul 23, 2026
a47577a
[None][infra] Waive 1 failed cases for main in pre-merge 49229 (#16786)
trtllm-agent Jul 23, 2026
e96c330
[https://nvbugs/6445456][fix] Restore inplace ops for functionalizati…
liji-nv Jul 23, 2026
96360f8
[https://nvbugs/6426850][test] Unwaive Qwen3.5 397B NVFP4 ADP4 TRTLLM…
liji-nv Jul 23, 2026
83733c4
[None][chore] Add NVTX ranges to per-iteration ADP sync points in PyE…
lancelly Jul 23, 2026
e16dcc5
[None][feat] Default GLM-5 to the Python KV-cache transceiver (#16524)
chuangz0 Jul 23, 2026
a1c6908
[None][feat] top-k: route decode to CuTe DSL GVR top-k in e2e (#16420)
limin2021 Jul 23, 2026
77774d9
[None][fix] Make FlashInfer sampling op wrappers opaque to Dynamo (#1…
qiaoxj07 Jul 23, 2026
f4e692d
[None][infra] Waive 4 failed cases for main in pre-merge 49550 (#16798)
trtllm-agent Jul 23, 2026
5a69240
[None][feat] Bind SourceIdentity to checkpoint artifacts (#16159)
chienchunhung Jul 23, 2026
526c1e3
[https://nvbugs/6448152][perf] make C++ context-transfer consensus as…
chienchunhung Jul 23, 2026
80b4eb3
[None][perf] Optimize fusedQKNormRope Kernel (#16633)
aswinvisva Jul 23, 2026
d95f665
[None][fix] Stage DSA indexer block_table H2D through pinned memory (…
Tabrizian Jul 23, 2026
35134c2
[None][infra] Unwaive test_proxy_fast_death tests (#16808)
dpitman-nvda Jul 24, 2026
e5287d6
[None][infra] Waive 2 failed cases for main in pre-merge 49489 (#16819)
trtllm-agent Jul 24, 2026
3c07ada
[TRTLLM-14474][chore] Remove legacy python relics and refresh docs af…
Wanli-Jiang Jul 24, 2026
7d3a4e9
[https://nvbugs/6426834][fix] Deflake test_kv_transfer: cap NIXL prog…
chuangz0 Jul 24, 2026
2801e94
[TRTLLM-14575][fix] MoE: fp32 accumulation in deferred MoEAllReduce f…
xwang233 Jul 24, 2026
16bdc6d
[None][feat] Support MARLIN MoE with MTP and attention DP + EP (#16597)
Wanli-Jiang Jul 24, 2026
37f59ac
[https://nvbugs/6485885][fix] Stop thinking-budget processor re-forci…
Wanli-Jiang Jul 24, 2026
0960d59
[#16767][fix] Fix DSpark rolling-window slot collision in disaggregat…
longlee0622 Jul 24, 2026
cf49225
[None][infra] Fix release check failure for .test_durations (#16784)
EmmaQiaoCh Jul 24, 2026
58b9d15
[None][perf] spec one-model sampling: greedy rows via top_k=1 instead…
Tabrizian Jul 24, 2026
2991994
[None][perf] Avoid implicit device-scalar syncs in DeepSeek-V4 ctx sp…
qiaoxj07 Jul 24, 2026
929f153
[TRTLLM-9920][feat] Add support for arbitrary KVCache transfer (#13055)
Tabrizian Jul 24, 2026
2fbb367
[https://nvbugs/6473161][chore] unwaive 4 TestLlama3_1 tests (#16824)
lori-ren Jul 24, 2026
8a6cb44
[None][test] Declare supported attention test phases (#16718)
yihwang-nv Jul 24, 2026
f1645c4
[TRTLLM-13948][test] Add DeepSeek R1/V3.2/V3-Lite disaggregated accur…
asfiyab-nvidia Jul 24, 2026
1fae43c
[TRTLLM-13349][perf] Fuse gemma RMSNorm into AllReduce for Qwen3-Next…
nv-guomingz Jul 24, 2026
c1f78d9
[None][perf] Fuse DeepSeek-V4 Indexer Q projection with CuTe DSL (#16…
mingyangHao Jul 24, 2026
afd0a66
[None][feat] Add support for MiniMax M3 and Qwen3.6 models in perform…
yufeiwu-nv Jul 24, 2026
2e2ed4e
[NVBUG-6379624][fix] Enable W4A8 checkpoint loading for Gemma4 K=V la…
Hudayday Jul 24, 2026
56dedb1
[None][infra] Waive 14 failed cases for main in post-merge 2855 (#16821)
trtllm-agent Jul 24, 2026
8514fa3
[None][feat] align time metrics between cpp and python cache transcei…
chuangz0 Jul 24, 2026
641bbf2
[TRTLLM-13948][test] Migrate DeepSeek R1/V3.2 disagg perf cases to tr…
nv-xtf Jul 24, 2026
75b39d4
[https://nvbugs/6482576][fix] Fall back to disagg_request_id in Pytho…
Shixiaowei02 Jul 24, 2026
a8b5409
[None][perf] Reordering torch.compile and cache-dit call (#16508)
BrianLi23 Jul 24, 2026
4d066d3
[None][fix] Do not treat a bare fmha_v2_cu directory as completed gen…
brnguyen2 Jul 24, 2026
2eef5fe
[TRTLLM-14512][feat] multiprocess disagg server prometheus client (#1…
reasonsolo Jul 24, 2026
021b435
[None][test] Enable gen_only + ctx_only DeepSeek-V4-Pro perf-sanity o…
chenfeiz0326 Jul 24, 2026
121cf56
[https://nvbugs/6457853][fix] Allow trtllm-gen MoE autotuner when loc…
dongfengy Jul 24, 2026
e6b7bd0
[https://nvbugs/6432953][fix] Fix MPI world heap corruption during te…
mikeiovine Jul 24, 2026
f5c2a07
[https://nvbugs/6020038][feat] Add NCCL-EP v0.1 MoE communication sup…
nv-lschneider Jul 24, 2026
9d7ef31
[None][fix] Fix GPT-OSS router token identity (#16760)
SimengLiu-nv Jul 24, 2026
4b7d719
[TRTLLM-14417][fix] Exclude ADP/cuda-graph dummy requests from specul…
xwang233 Jul 24, 2026
b8e4594
[https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gp…
xxi-nv Jul 25, 2026
bfb0fca
[https://nvbugs/6463822][fix] Fix LTX2 CUDA graph test leak issue (#1…
yibinl-nvidia Jul 25, 2026
cf44a1c
[https://nvbugs/6465993][fix] use attention cache dtype for disaggreg…
chienchunhung Jul 25, 2026
1562a07
[None][feat] Support DeepSeek-V4 in layer_wise_benchmarks (#16774)
ruodil Jul 26, 2026
9de6c94
[None][perf] Skip DeepGEMM clean_logits in DSA indexer prefill on cus…
dc3671 Jul 27, 2026
9129f4e
[None][infra] Auto-update test durations from OpenSearch (last 7 days)
tensorrt-cicd Jul 27, 2026
08289b6
[None][perf] Optimize Blackwell fused MHC half-MMA kernel (#16799)
MengmSun Jul 27, 2026
b8ff548
[None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalli…
hyukn Jul 27, 2026
1ae9b86
[None][infra] Waive 21 failed cases for main in post-merge 2862 (#16882)
trtllm-agent Jul 27, 2026
d12c85e
[https://nvbugs/6507109][infra] Split slow DGX B300 attention unit te…
yuxianq Jul 27, 2026
da39470
[https://nvbugs/6479324][test] Remove waiver for fixed qwen3_5_4b_fp8…
VALLIS-NERIA Jul 27, 2026
55e9b2f
[None][fix] Resolve NVFP4 mixed-precision base layers for the DSpark …
tianyuz-nv Jul 27, 2026
49e16c9
[https://nvbugs/6433376][fix] Update the Dense test to mirror the MoE…
trtllm-agent Jul 27, 2026
b91ffda
[TRTLLM-13642][feat] Add perf sanity tests for Llama-3.1-8B and Gemma…
moraxu Jul 27, 2026
7982aa9
[https://nvbugs/6501376][fix] Test-only fix — drop the `if hidden_siz…
trtllm-agent Jul 27, 2026
1dd8b97
[None][feat] Add kimi_k2/glm_5 grouped routing and fused router to be…
guqiqi Jul 27, 2026
151db5d
[https://nvbugs/6157892][fix] Mistral format refactor (#15123)
evezhier Jul 27, 2026
9f5b377
[None][test] Adjust timeout cases in QA perf test (#16894)
yufeiwu-nv Jul 27, 2026
aae253e
[#15673][fix] Enable CUDA core fast path for SM89/SM120/SM121 (#12705)
mihai-chiorean Jul 27, 2026
1b4ffc0
[TRTLLM-14475][chore] Self-sufficient transfer-agent dlopen and drop …
Wanli-Jiang Jul 27, 2026
0f5e15e
[None][fix] Clarify explicit post-merge stage CI label (#16886)
yibinl-nvidia Jul 27, 2026
e95cb90
[None][fix] Drop stale benchmarks copies from Dockerfile.multi (#16884)
jieli-matrix Jul 27, 2026
6046f34
[https://nvbugs/6198785][fix] Unify phase-1 CUDA graph cleanup (#16763)
Mgluhovskoi Jul 27, 2026
888fa17
[TRTLLMINF-40][fix] Introduce a SLURM dispatcher pod "finalizer" (#16…
dpitman-nvda Jul 27, 2026
155847d
[https://nvbugs/6507081][fix] Refresh the fakes only — add `reasoning…
trtllm-agent Jul 27, 2026
9392453
[None][fix] Fix disaggregated draft token accounting (#16805)
SimengLiu-nv Jul 27, 2026
9fe5853
[None][feat] VisualGen TP with Attn2D (#16677)
belgarten-nv Jul 27, 2026
cfeca00
[None][feat] Generic Mixed Modality Support (#16337)
aswinvisva Jul 27, 2026
d7be54b
[None][perf] Enable FLUX2 VisualGen fused NVFP4 SwiGLU path (#16143)
pst2154 Jul 28, 2026
c4f3353
[None][feat] Support MiniCPM-V 4.6 (image + video) on the PyTorch bac…
hNSBQZ Jul 28, 2026
798e419
[https://nvbugs/6479837][fix] Fix OOM of Qwen3_5_35B on a single a100…
JadoTu Jul 28, 2026
89c6635
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jul 28, 2026
f3d4c85
[None][perf] GVR top-K decode: enable R0 histogram-ladder admission b…
longcheng-nv Jul 28, 2026
7a64f26
[https://nvbugs/6479863][fix] Use scalar SwiGLU limit for DeepSeek V4…
lfr-0531 Jul 28, 2026
ede2cad
[TRTLLM-14609][chore] Remove legacy MoE path in CuteDslFusedMoE (#16863)
xxi-nv Jul 28, 2026
343a20b
[TRTLLM-14609][chore] Remove legacy MoE path in DeepGemmFusedMoE (#16…
xxi-nv Jul 28, 2026
cfebf19
[None][feat] Add Qwen-Image-Layered baseline support (#15096)
yumin066 Jul 28, 2026
5e4a154
[None][test] Enable session prefetch for all test stages (#16770)
sunnyqgg Jul 28, 2026
729eb4e
[None][test] Add missing test durations for MiniMaxM3, Step3_7, and M…
xinhe-nv Jul 28, 2026
99d61db
[None][test] Waive 7 failed cases for main in QA CI (#16934)
trtllm-agent Jul 28, 2026
53659b1
[None][test] Fix test_perf.py to accept kv cache manager v2 format (#…
yufeiwu-nv Jul 28, 2026
becf773
[None][fix] Fix nemotron-h quant and loading config (#16833)
Wanli-Jiang Jul 28, 2026
6219c2e
[https://nvbugs/6450333][test] Unwaive DeepSeek V4 Flash auto dtype t…
lfr-0531 Jul 28, 2026
65b1e53
[https://nvbugs/6435112][test] Unwaive Wan 2.2 I2V perf sanity test (…
taianz-nv Jul 28, 2026
d6a2d25
[TRTLLM-11875][feat] BREAKING: MambaCacheManager based on KVCacheMana…
VALLIS-NERIA Jul 28, 2026
2a7231d
[None][test] Update CODEOWNERS to refine QA ownership by adding speci…
yufeiwu-nv Jul 28, 2026
201adec
[https://nvbugs/6510284][fix] Cap gen-only benchmark queue size (#16915)
chienchunhung Jul 28, 2026
e552a61
[TRTLLM-14502][feat] LTX-2 two-stage: dual-topology parallel Stage 2 …
luyiyun1021 Jul 28, 2026
03d4be4
[https://nvbugs/6240584][fix] Qwen3ToolParser: bare-JSON fallback for…
JunyiXu-nv Jul 28, 2026
5ac2259
[None][feat] Batched physical KV-cache compaction for KV cache compre…
Hudayday Jul 28, 2026
90f7868
[None][test] Waive 1 failed cases for main in QA CI (#16947)
trtllm-agent Jul 28, 2026
c82ae76
[https://nvbugs/6503293][fix] Restore whole-node GPU visibility for d…
JacobHu-NV Jul 28, 2026
4d4e8aa
[None][fix] Update DeepSeek V4 Flash-Base MoE backend configuration i…
yufeiwu-nv Jul 28, 2026
1b9cbfa
[TRTLLM-14609][chore] Remove legacy MoE path in CutlassFusedMoE (#16861)
xxi-nv Jul 28, 2026
e2b7145
[None][fix] Keep chunked-MoE size vectors identical across attention-…
dongfengy Jul 28, 2026
1f1acea
[TRTLLM-13233][feat] Support no_repeat_ngram_size in TorchSampler and…
zhaoyangwang-nvidia Jul 28, 2026
78dd19c
[TRTLLM-11780][feat] Wan 2.2 layernorm + shiftscale + quant fusion (#…
o-stoner Jul 28, 2026
9833c64
[None][test] Waive 7 failed cases for main in QA CI (#16950)
trtllm-agent Jul 28, 2026
a63ef1f
[None][test] Stabilize scaffolding OpenAI worker tests (#16853)
Mgluhovskoi Jul 28, 2026
095ab2c
[None][fix] Don't re-run the SLURM monitor on a terminal job failure …
dpitman-nvda Jul 28, 2026
05edf29
[https://nvbugs/6484986][fix] cancel pending UCX receive (#16688)
chienchunhung Jul 28, 2026
3f70158
[https://nvbugs/6507955][fix] use net_max_seq_len for request admisso…
bo-nv Jul 28, 2026
feb83ec
[https://nvbugs/6226016][fix] Avoid trusting request-controlled route…
yibinl-nvidia Jul 28, 2026
f9ac468
[TRTLLM-12341][feat] Add Whisper support to the PyTorch backend (#16141)
pranav-nvidia Jul 28, 2026
4235bef
[TRTLLM-14135][feat] Add Qwen-Image-Edit-2511 support (#16095)
yibinl-nvidia Jul 29, 2026
cb44a40
[TRTLLM-14609][chore] Remove legacy MoE path in TRTLLMGenFusedMoE (#1…
xxi-nv Jul 29, 2026
f6125fb
[None][perf] Avoid Index-K cache materialization for MSA (#16856)
peihu-nv Jul 29, 2026
0f542f3
[None][chore] add user (#16754)
tburt-nv Jul 29, 2026
058bbe3
[None][test] Waive 4 failed cases for main in QA CI (#16979)
trtllm-agent Jul 29, 2026
7e8eb8f
[None][feat] Update CuTeDSL MegaMoE kernels (#16190)
Barry-Delaney Jul 29, 2026
71fbbc2
[None][test] Waive 1 failed cases for main in QA CI (#16983)
trtllm-agent Jul 29, 2026
c4afe88
[TRTLLM-14609][chore] Remove legacy MoE path in DenseGEMMFusedMoE (#1…
xxi-nv Jul 29, 2026
f5cbe6b
[None][feat] Improve cute dsl radix top-k (#15756)
limin2021 Jul 29, 2026
15dbae9
[None][test] Waive 7 failed cases for main in QA CI (#16982)
trtllm-agent Jul 29, 2026
87afac8
[#15327][feat] Add per-request priority support to OpenAI chat/comple…
sopwg612 Jul 29, 2026
6b75514
[None][test] Waive 1 failed cases for main in QA CI (#16984)
trtllm-agent Jul 29, 2026
0f3f850
[None][test] Waive 1 failed cases for main in QA CI (#16985)
trtllm-agent Jul 29, 2026
d9c53cc
[None][infra] Waive 1 failed cases for main in pre-merge 50457 (#16978)
trtllm-agent Jul 29, 2026
3449011
[None][infra] Waive 5 failed cases for main in post-merge 2865 (#16989)
trtllm-agent Jul 29, 2026
ebd197f
[TRTLLM-13409][test] fail fast + surface server logs when a perf-sani…
JunyiXu-nv Jul 29, 2026
9d508eb
[TRTLLM-14541][fix] VisualGen: deterministic autotuner tactics across…
luyiyun1021 Jul 29, 2026
097cbc1
[https://nvbugs/6503299][fix] Default fabric memory KV pool for Pytho…
chuangz0 Jul 29, 2026
a1e5771
[None][perf] Preserve default V2 KV cache pool sizing (#16783)
2ez4bz Jul 29, 2026
2b2bf97
[None][fix] enable static EPLB for the one-model DSpark drafter (#16938)
longlee0622 Jul 29, 2026
09d8715
[https://nvbugs/6255417][fix] Unwaive qwen3next ci test (#16924)
JadoTu Jul 29, 2026
b883d55
[None][feat] cache transceiver test in Perf sanity (#16674)
chuangz0 Jul 29, 2026
c9a2629
[https://nvbugs/6490033][fix] Relaxed the assertion to accept both `i…
trtllm-agent Jul 29, 2026
4b2e48b
[None][fix] Keep MRoPE delta read slots dense across mixed batches (#…
yechank-nvidia Jul 29, 2026
2f113fd
[https://nvbugs/6210714][test] Unwaive TestQwen3_5_35B_A3B fp8 block …
VALLIS-NERIA Jul 29, 2026
38389f9
[https://nvbugs/6465993][fix] unwaive mamba tests (#16941)
bo-nv Jul 29, 2026
e240d4a
[TRTLLM-14571][infra] Enable container-local AutoTuner cache in CI (#…
YihuiLu512 Jul 29, 2026
2341c70
[None][fix] Fix Qwen3.5 weight-load memory growth and MTP CUTLASS fal…
Wanli-Jiang Jul 29, 2026
af64dff
[TRTLLM-14551][perf] avoid GDN state reset host synchronization (#16716)
liji-nv Jul 29, 2026
99bdffc
[https://nvbugs/6487040][test] Wait for gen-log end-of-write sentinel…
chenfeiz0326 Jul 29, 2026
f20ea65
[https://nvbugs/6337224][fix] Update PERF_SANITY_DIR to include `aggr…
tensorrt-cicd Jul 29, 2026
b004352
[TRTLLM-14736][chore] Split the sampler package into per-feature modu…
zhaoyangwang-nvidia Jul 29, 2026
2e1a997
[#8384][fix] use dict.get() instead of getattr() for rope_scaling dic…
wojciech-wais Jul 29, 2026
fa5be69
[None][feat] Multimodal encoder cache: per-item partial hits (#16817)
aswinvisva Jul 29, 2026
c45ad83
[https://nvbugs/6163690][fix] Use PreTrainedTokenizerFast in trtllm-b…
pamelap-nvidia Jul 29, 2026
4b9012b
[None][perf] Fuse index-q/index-k projections in MinimaxM3 (#16904)
brb-nv Jul 29, 2026
960530b
[None][perf] Fuse MiniMax-M3 MoE routing (#16859)
peihu-nv Jul 29, 2026
9c345f8
[https://nvbugs/6529626][fix] Pin mcp<2.0.0 and unwaive the scaffoldi…
JunyiXu-nv Jul 30, 2026
86e7571
[None][perf] Add Qwen Image VisualGen perf fastpaths (#16142)
pst2154 Jul 30, 2026
d0543dc
[None][fix] Increase max top logprobs limit (#16851)
yibinl-nvidia Jul 30, 2026
8624ec9
[https://nvbugs/6424956][fix] Support large FP8 quantization grids (#…
lfr-0531 Jul 30, 2026
d0b60a1
[https://nvbugs/6523880][fix] Restored the `legacy-files.txt` entry a…
trtllm-agent Jul 30, 2026
ffba1a6
[None][chore] update DeepGEMM to 2.6.1 (#16673)
Barry-Delaney Jul 30, 2026
c0eac26
[None][infra] Waive 15 failed cases for main in post-merge 2869 (#17041)
trtllm-agent Jul 30, 2026
c0df3b3
[None][test] update coderabbit prompt (#16988)
xinhe-nv Jul 30, 2026
1c052b4
[None][fix] cascade: wire workspace regardless of multi_block_mode (#…
Nic-bit Jul 30, 2026
49efbc1
[None][infra] Remove unused GitLab token env from pytest (#16946)
mzweilz Jul 30, 2026
78ab4e7
[None][feat] Support the DMD2-distilled Cosmos3-Super-Text2Image-4Ste…
ishovkun Jul 30, 2026
60fddd8
[None][feat] MTP one-model `advanced_sampling_mode`: skip redundant t…
jhaotingc Jul 30, 2026
af1d452
[None][test] Waive 1 failed cases for main in QA CI (#17045)
trtllm-agent Jul 30, 2026
787ee94
[None][test] Waive 1 failed cases for main in QA CI (#17044)
trtllm-agent Jul 30, 2026
eed4d5a
[https://nvbugs/6517842][fix] Handle mutable tensor lists in remove c…
liji-nv Jul 30, 2026
6eb951e
[https://nvbugs/6276841][fix] When torch_compile=True, pass kv_cache_…
tensorrt-cicd Jul 30, 2026
da5b62f
[TRTLLMINF-250][infra] Change jnlp image from urm.nvidia.com to artif…
yiqingy0 Jul 30, 2026
05efb7d
[None][infra] Select UCX env for perf sanity by cluster name (#16725)
chuangz0 Jul 30, 2026
2a5baab
[None][fix] SA spec dec: promote accepted hybrid recurrent states in-…
brnguyen2 Jul 30, 2026
88bfcae
[None][infra] Container vulnerability fix (#16694)
yuanjingx87 Jul 30, 2026
a146f66
[None][perf] Fused SwiGLU-OAI (#16905)
brb-nv Jul 30, 2026
a9dcdb3
[None][fix] Add mutex to avoid potentially concurrent modifications t…
yihwang-nv Jul 30, 2026
01a04f0
[TRTLLM-13229][feat] implement repetition / frequency / presence pena…
lori-ren Jul 30, 2026
83c0eb1
[https://nvbugs/6463829][fix] Fix fp8 MoE test (#16514)
brb-nv Jul 30, 2026
b8a4af1
[https://nvbugs/6487836][chore] unwaive test_performance_alignment[1]…
tburt-nv Jul 30, 2026
1b0b2ae
[None][test] Stabilize scaffolding LLM tests (#16966)
Mgluhovskoi Jul 30, 2026
5b6a3e5
[None][test] Stabilize TRTLLM scaffolding worker test (#16964)
Mgluhovskoi Jul 30, 2026
353a4ee
[None][fix] SpecDecOneEngineForCausalLM: accept optional hidden_size/…
brnguyen2 Jul 30, 2026
c083da6
[TRTLLM-14287][feat] Qwen Image CFG parallelism support (#16384)
yibinl-nvidia Jul 30, 2026
7f7dccf
[None][feat] KVCacheManagerV2 C++ translation (#14047)
lowsfer Jul 30, 2026
5c5ef98
[https://nvbugs/6305404][chore] Unwaive DeepSeek V3 Lite L0 test (#16…
Mgluhovskoi Jul 30, 2026
9e6af13
[None][infra] Waive 1 failed cases for main in pre-merge 50924 (#17079)
trtllm-agent Jul 31, 2026
85620fd
[TRTLLM-13230][feat] support min_p sampling for TorchSampler (#16590)
lori-ren Jul 31, 2026
138eb43
[TRTLLM-14609][chore] Remove ENABLE_CONFIGURABLE_MOE escape hatch and…
xxi-nv Jul 31, 2026
8e602fa
[None][refactor] Clean up model paths and remove deprecated configura…
yufeiwu-nv Jul 31, 2026
a2df74e
[None][feat] Enable MM encoder cache on Qwen3.x and Gemma4 VLMs (#16662)
2ez4bz Jul 31, 2026
10a9432
[TRTLLM-14304][feat] Integrate embeddings cache with encoder side-str…
2ez4bz Jul 31, 2026
dc7f325
[https://nvbugs/6523767][fix] Size MPI worker-identity barrier timeou…
trtllm-agent Jul 31, 2026
2fa17cb
[https://nvbugs/6287561][fix] Add `get_sm_version() < 90` check at th…
tensorrt-cicd Jul 31, 2026
be9afad
[None][infra] Add blossom-ci authorized users (#17109)
yiqingy0 Jul 31, 2026
b118fc3
[TRTLLM-14779][fix] Clear capture-only sampling override from cached …
xwang233 Jul 31, 2026
9f2d3c6
[None][feat] Log running metric estimates during long lm-eval runs (#…
brnguyen2 Jul 31, 2026
f10a22e
[None][fix] Fix Qwen3Next MoE expert-quant probe and GDN verify tenso…
Wanli-Jiang Jul 31, 2026
d924d9f
[TRTLLM-14010][feat] report KV cache transfer state on executor hangs…
bo-nv Jul 31, 2026
19bdfd3
[TRTLLM-14511][feat] BREAKING: refactor per-request perf metrics for …
reasonsolo Jul 31, 2026
e6568c3
[None][infra] Bump version to 1.3.0rc24 (#17070)
mikeiovine Jul 31, 2026
932d3b9
[TRTLLM-14177][feat] support reference images in FLUX.2 (#16644)
karljang Jul 31, 2026
72434f8
[TRTLLM-14709][infra] Require packaging>=24.2 for FlashInfer source b…
brnguyen2 Jul 31, 2026
d91d41a
[https://nvbugs/6451425][fix] Remove llama3 eagle test waive (#17128)
mikeiovine Jul 31, 2026
574265e
[None][perf] Kimi K3: combine MR99+MR100+MR102 on latest feat (with M…
litaotju Jul 30, 2026
2e4c343
[None][fix] Kimi K3: unblock !103 - deferred-finalize scales fix, fus…
moraxu Jul 30, 2026
066bab1
[None][feat] Support the DMD2-distilled Cosmos3 4-step image-to-video…
ishovkun Jul 31, 2026
575d7f3
[None][chore] Apply pre-commit formatting to the cherry-picked changes
moraxu Jul 30, 2026
7003910
[None][chore] Move FORCE_SEPARATED_ROUTING below the imports to fix r…
moraxu Jul 31, 2026
55d55ff
[None][perf] AllReduce + ResidualAdd + RMSNorm (#17091)
brb-nv Jul 31, 2026
b1b8dbd
[None][fix] Bind explicit DP rank for new conversations (#16815)
krishung5 Jul 31, 2026
e34d3d4
[https://nvbugs/6537081][fix] Fix import error (#17112)
2ez4bz Jul 31, 2026
06ec703
Merge branch 'main' into feat/kimi_k3
brnguyen2 Jul 31, 2026
46b0d0d
Merge remote-tracking branch 'origin/feat/kimi_k3' into merge/main-in…
brnguyen2 Jul 31, 2026
3328157
[TRTLLM-14810][fix] Use tensorrt_llm::DataType in MNNVL allreduce MoE…
brnguyen2 Jul 31, 2026
78ae85b
Merge commit 'refs/tmp/reconcile' into merge/main-into-kimi-k3
brnguyen2 Jul 31, 2026
ff30e9f
[TRTLLM-14810][fix] Pass MLA multi-CTAS KV counter buffer only on the…
brnguyen2 Jul 31, 2026
fc38923
[TRTLLM-14810][fix] Post-merge fixups: formatting, reuse snapshot int…
brnguyen2 Aug 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 1 addition & 3 deletions .claude/skills/exec-local-compile/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ git checkout main && git pull
Run the build command (**incremental by default** — omit `-c`/`--clean` unless explicitly requested or the incremental build fails):

```bash
./scripts/build_wheel.py --trt_root /usr/local/tensorrt --benchmarks --use_ccache -a "<arch>" -f --nvtx
./scripts/build_wheel.py --use_ccache -a "<arch>" -f --nvtx
```

Replace `<arch>` with the target GPU architecture (see Architecture Reference below). If not specified by the user, auto-detect from `nvidia-smi`.
Expand All @@ -66,8 +66,6 @@ python3 -c "import tensorrt_llm; print(tensorrt_llm.__version__)"

| Flag | Description |
|------|-------------|
| `--trt_root /usr/local/tensorrt` | TensorRT installation path (standard in NVIDIA containers) |
| `--benchmarks` | Build the C++ benchmarks |
| `-a "<arch>"` | Target GPU architecture(s) |
| `--nvtx` | Enable NVTX markers for profiling |
| `--use_ccache` | Use ccache for faster recompilation |
Expand Down
3 changes: 0 additions & 3 deletions .claude/skills/exec-slurm-compile/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -204,8 +204,6 @@ A successful build ends with a message like `Successfully built tensorrt_llm` or

| Flag | Description |
|------|-------------|
| `--trt_root /usr/local/tensorrt` | TensorRT installation path (standard in NVIDIA containers) |
| `--benchmarks` | Build the C++ benchmarks |
| `-a "100-real"` | Target architecture — `100` for Blackwell, `90` for Hopper, etc. |
| `--nvtx` | Enable NVTX markers for profiling |
| `--no-venv` | Skip virtual environment creation |
Expand All @@ -228,7 +226,6 @@ Common architecture values:
| `sbatch: error: invalid partition` | Verify partition name with `sinfo -s` |
| `sbatch: error: invalid account` | Check available accounts with `sacctmgr show assoc user=$USER` |
| Container image not found | Verify the `.sqsh` path exists and is readable |
| Build fails with missing TensorRT | Ensure `--trt_root` points to the correct path inside the container |
| Build OOM (out of memory) | Reduce parallelism with `-j <N>` flag to `build_wheel.py` |
| `srun: error: Unable to create step` | The node may lack enroot/pyxis — check with cluster admin |
| Job stuck in `PD` state | Check `squeue -j <id> -o %R` for the reason (e.g., resource limits, priority) |
Expand Down
4 changes: 1 addition & 3 deletions .claude/skills/exec-slurm-compile/scripts/compile.sh
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
# Usage: compile.sh <repo_dir> [build_wheel_args...]
#
# Default build_wheel.py flags:
# --trt_root /usr/local/tensorrt --benchmarks -a "100-real" --nvtx --no-venv
# -a "100-real" --nvtx --no-venv
# Any extra arguments after repo_dir are forwarded to build_wheel.py,
# overriding the defaults above.

Expand All @@ -36,8 +36,6 @@ if [[ $# -gt 0 ]]; then
else
echo "[compile.sh] Running default build command"
python3 ./scripts/build_wheel.py \
--trt_root /usr/local/tensorrt \
--benchmarks \
-a "100-real" \
--nvtx
fi
39 changes: 35 additions & 4 deletions .coderabbit.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,26 @@ reviews:
auto_title_placeholder: '@coderabbitai title'
auto_title_instructions: 'Format: "[<category>] <title>". Category must be one of: fix, feat, doc, infra, style, refactor, perf, test, chore, revert. Enclose the category in square brackets. Title should be concise (<= 60 chars). Example: "[feat] Add logit_bias support".'
commit_status: false
collapse_walkthrough: true
high_level_summary_instructions: |
Always produce two review sections in the summary:

**Dev Engineer Review**
Review all changes for correctness and consistency, including:
- Code changes: correctness, performance, API consistency (CODING_GUIDELINES.md), error handling, regressions.
- Config files: valid values, no typos, consistency with related configs, no unintended scope changes.
- Test list files (test-db/, qa/, waives.txt): correct format, valid test paths, appropriate bug references, no duplicates.

**QA Engineer Review**
Always include this section when any files under tests/ are touched.
For test-list-only changes (only tests/integration/test_lists/ files):
- List which test-db/ or qa/ files were modified and what entries were added or removed.
- Verdict: "needs follow-up" if CBTS coverage data is unavailable, otherwise "sufficient" or "insufficient".
For test-code changes (files outside tests/integration/test_lists/):
- List test functions added, modified, or removed.
- State whether each is covered in tests/integration/test_lists/ (test-db/ for CI, qa/ for manual QA).
- Verdict: sufficient, insufficient, or needs follow-up.
If no test files are touched, write "No test changes."
collapse_walkthrough: false
assess_linked_issues: true
related_issues: true
related_prs: true
Expand All @@ -31,15 +50,27 @@ reviews:
poem: false
review_status: false
auto_review:
auto_incremental_review: false
auto_incremental_review: true
drafts: false
base_branches: ["main", "release/.+"]
path_instructions:
- path: "tests/**"
instructions: |
Act as a QA engineer reviewing test changes and coverage for TensorRT-LLM.
Keep feedback actionable: suggest concrete list file names and whether
coverage is sufficient, insufficient, or needs follow-up outside the PR.
Always produce a test coverage summary, even if no issues are found.

If the change touches ONLY files under tests/integration/test_lists/ (no test-code changes):
- Report which test-db/ or qa/ list files were modified and what entries were added or removed.
- Do NOT require changed test functions for this path.
- Use verdict "needs follow-up" when cbts_touchmap.sqlite or a CBTS coverage report is unavailable
to confirm the impacted test scope; otherwise use "sufficient" or "insufficient".

If the change includes test-code files (outside tests/integration/test_lists/), the summary must include:
1. Which test functions were added, modified, or removed.
2. Whether each changed test is listed in the appropriate test list files under
tests/integration/test_lists/ (test-db/ for CI, qa/ for manual QA).
3. A coverage verdict: sufficient, insufficient, or needs follow-up.
Keep feedback actionable: reference concrete list file names when suggesting additions.
- path: "tests/integration/test_lists/qa/**"
instructions: |
Files here are manually-triggered QA perf/regression lists, maintained
Expand Down
3 changes: 2 additions & 1 deletion .gitattributes
Original file line number Diff line number Diff line change
Expand Up @@ -16,4 +16,5 @@ docs/source/blogs/media/tech_blog10_full_strategy_performance.png filter=lfs dif
docs/source/blogs/media/tech_blog10_context_wait_performance.png filter=lfs diff=lfs merge=lfs -text
cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/kernelMetaInfo_cubin.cpp filter=lfs diff=lfs merge=lfs -text
cpp/tensorrt_llm/kernels/decoderMaskedMultiheadAttention/cubin/xqa_kernel_cubin.cpp filter=lfs diff=lfs merge=lfs -text
tensorrt_llm/_torch/visual_gen/cute_dsl_kernels/blackwell/attention/cubins/*/*/*.so filter=lfs diff=lfs merge=lfs -text
docs/source/blogs/media/tech_blog26_deepseek_v4_hybrid_attention.png filter=lfs diff=lfs merge=lfs -text
docs/source/blogs/media/tech_blog26_deepseek_v4_mhc_moe.png filter=lfs diff=lfs merge=lfs -text
24 changes: 12 additions & 12 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
# Infra @NVIDIA/trt-llm-infra-devs
# Agent config @NVIDIA/trt-llm-agent-devs
# Docs / Examples @NVIDIA/trt-llm-doc-owners
# QA @NVIDIA/trt-llm-qa
# QA @NVIDIA/trt-llm-qa / qa-perf / qa-function
# Runtime @NVIDIA/trt-llm-runtime-devs
# Kernels - Misc @NVIDIA/trt-llm-kernels-devs
# Models @NVIDIA/trt-llm-models-devs
Expand All @@ -31,14 +31,12 @@
/tensorrt_llm/commands/eval.py @NVIDIA/trt-llm-devs
/tensorrt_llm/evaluate @NVIDIA/trt-llm-devs
/tensorrt_llm/tools @NVIDIA/trt-llm-devs
/tests/integration/test_lists/test-db @NVIDIA/trt-llm-devs @NVIDIA/trt-llm-qa @NVIDIA/trt-llm-infra-devs
/tests/integration/test_lists/waives.txt @NVIDIA/trt-llm-devs @NVIDIA/trt-llm-qa @NVIDIA/trt-llm-infra-devs
/tests/integration/test_lists/test-db @NVIDIA/trt-llm-devs @NVIDIA/trt-llm-qa-function @NVIDIA/trt-llm-infra-devs
/tests/integration/test_lists/waives.txt @NVIDIA/trt-llm-devs @NVIDIA/trt-llm-qa-function @NVIDIA/trt-llm-infra-devs
/tests/test_common @NVIDIA/trt-llm-devs
/tests/unittest @NVIDIA/trt-llm-devs

# ===== TensorRT backend (will be deprecated soon) — also on the trt-llm-devs fallback =====
/cpp/include/tensorrt_llm/plugins @NVIDIA/trt-llm-devs
/cpp/tensorrt_llm/plugins @NVIDIA/trt-llm-devs
/tensorrt_llm/builder.py @NVIDIA/trt-llm-devs
/tensorrt_llm/commands/build.py @NVIDIA/trt-llm-devs
/tensorrt_llm/commands/prune.py @NVIDIA/trt-llm-devs
Expand Down Expand Up @@ -95,6 +93,8 @@
# ===== QA =====
/tests/integration/defs @NVIDIA/trt-llm-devs @NVIDIA/trt-llm-qa @NVIDIA/trt-llm-infra-devs
/tests/integration/test_lists/qa @NVIDIA/trt-llm-qa
/tests/integration/test_lists/qa/llm_perf_* @NVIDIA/trt-llm-qa-perf
/tests/integration/test_lists/qa/llm_function_* @NVIDIA/trt-llm-qa-function

# ===== RUNTIME =====
/cpp/include/tensorrt_llm/batch_manager @NVIDIA/trt-llm-runtime-devs
Expand All @@ -105,7 +105,6 @@
/cpp/tensorrt_llm/batch_manager @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/common @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/executor @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/executor_worker @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/layers @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/nanobind @NVIDIA/trt-llm-runtime-devs
/cpp/tensorrt_llm/runtime @NVIDIA/trt-llm-runtime-devs
Expand Down Expand Up @@ -169,13 +168,12 @@
/examples/llm-api/quickstart_multimodal.py @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-doc-owners
/examples/models @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-doc-owners
/examples/serve/*multimodal* @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-doc-owners
/scripts/build_cpp_examples.py @NVIDIA/trt-llm-models-devs
/scripts/generate_config_database_tests.py @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-doc-owners
/scripts/generate_config_table.py @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-doc-owners
/tensorrt_llm/_torch/models @NVIDIA/trt-llm-models-devs
/tensorrt_llm/_torch/modules/mamba @NVIDIA/trt-llm-models-devs
/tensorrt_llm/quantization @NVIDIA/trt-llm-models-devs
/tests/integration/defs/accuracy/test_llm_api_pytorch_multimodal.py @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-qa
/tests/integration/defs/accuracy/test_llm_api_pytorch_multimodal.py @NVIDIA/trt-llm-models-devs @NVIDIA/trt-llm-qa-function
/tests/unittest/_torch/modeling @NVIDIA/trt-llm-models-devs
/tests/unittest/_torch/models @NVIDIA/trt-llm-models-devs
/tests/unittest/_torch/modules/mamba @NVIDIA/trt-llm-models-devs
Expand Down Expand Up @@ -212,6 +210,7 @@
/cpp/tensorrt_llm/batch_manager/blockKey* @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tensorrt_llm/batch_manager/evictionPolicy* @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tensorrt_llm/batch_manager/kvCache* @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2 @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tensorrt_llm/nanobind/batch_manager/kvCacheManager* @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tests/unit_tests/batch_manager/blockKey* @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tests/unit_tests/batch_manager/evictionPolicy* @NVIDIA/trt-llm-kv-cache-manager-devs
Expand Down Expand Up @@ -246,9 +245,9 @@
/tensorrt_llm/disaggregated_params.py @NVIDIA/trt-llm-disagg-devs
/tensorrt_llm/serve/openai_disagg_server.py @NVIDIA/trt-llm-disagg-devs
# Disagg tests: co-own with the owning team so disagg-devs review disagg-test changes.
/tests/integration/defs/accuracy/*disagg* @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa
/tests/integration/defs/disaggregated @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa
/tests/integration/defs/stress_test/disagg_cancel @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa
/tests/integration/defs/accuracy/*disagg* @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa-serving
/tests/integration/defs/disaggregated @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa-serving
/tests/integration/defs/stress_test/disagg_cancel @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-qa-serving
/tests/scripts/perf-sanity/disaggregated @NVIDIA/trt-llm-perf-devs @NVIDIA/trt-llm-disagg-devs
/tests/scripts/perf/disaggregated @NVIDIA/trt-llm-perf-devs @NVIDIA/trt-llm-disagg-devs
/tests/unittest/_torch/executor/*disagg* @NVIDIA/trt-llm-runtime-devs @NVIDIA/trt-llm-disagg-devs
Expand Down Expand Up @@ -404,7 +403,7 @@
/scripts/check_auto_deploy_imports.py @NVIDIA/trt-llm-torch-autodeploy-devs
/scripts/check_model_registry.py @NVIDIA/trt-llm-torch-autodeploy-devs
/tensorrt_llm/_torch/auto_deploy @NVIDIA/trt-llm-torch-autodeploy-devs
/tests/integration/defs/accuracy/test_llm_api_autodeploy.py @NVIDIA/trt-llm-torch-autodeploy-devs @NVIDIA/trt-llm-qa
/tests/integration/defs/accuracy/test_llm_api_autodeploy.py @NVIDIA/trt-llm-torch-autodeploy-devs @NVIDIA/trt-llm-qa-function
/tests/unittest/_torch/auto_deploy @NVIDIA/trt-llm-torch-autodeploy-devs
/tests/unittest/auto_deploy @NVIDIA/trt-llm-torch-autodeploy-devs
/tests/integration/defs/accuracy/test_llm_api_autodeploy.py @NVIDIA/trt-llm-torch-autodeploy-devs @NVIDIA/trt-llm-qa-function
Expand Down Expand Up @@ -441,6 +440,7 @@
/cpp/tensorrt_llm/batch_manager/allocateKvCache.cpp @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tests/unit_tests/batch_manager/kvCacheManagerTest.cpp @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tests/unit_tests/batch_manager/kvCacheUtilsTest.cpp @NVIDIA/trt-llm-kv-cache-manager-devs
/tensorrt_llm/_torch/attention_backend/sparse/*/cache_manager.py @NVIDIA/trt-llm-kv-cache-manager-devs
/tensorrt_llm/_torch/pyexecutor/kv_cache_manager_v2.py @NVIDIA/trt-llm-kv-cache-manager-devs
/tensorrt_llm/_torch/pyexecutor/resource_manager.py @NVIDIA/trt-llm-kv-cache-manager-devs
/cpp/tensorrt_llm/nanobind/batch_manager/kvCacheManager.h @NVIDIA/trt-llm-kv-cache-manager-devs
Expand Down
8 changes: 8 additions & 0 deletions .github/workflows/blossom-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,9 @@ jobs:
"achartier",
"ajrasane",
"alec-flowers",
"AlessioNetti",
"alexmsettle",
"allisonlim-nv",
"ameynaik-hub",
"amirkl94",
"amitz-nv",
Expand Down Expand Up @@ -94,6 +96,7 @@ jobs:
"chzblych",
"cjluo-nv",
"crazydemo",
"daichu-nv",
"DanBlanaru",
"danielafrimi",
"davidclark-nv",
Expand Down Expand Up @@ -187,6 +190,7 @@ jobs:
"JunyiXu-nv",
"JyChang012",
"kaiyux",
"Kambili",
"kanghui0204",
"karljang",
"karthikvetrivel",
Expand Down Expand Up @@ -228,6 +232,7 @@ jobs:
"MatthiasKohl",
"mayani-nv",
"meenchen",
"MengmSun",
"mgluhovskoi",
"mikeiovine",
"milesial",
Expand Down Expand Up @@ -276,6 +281,7 @@ jobs:
"PerkzZheng",
"poweiw",
"pranav-nvidia",
"pst2154",
"qiangxu1996",
"qiaoxj07",
"QiJune",
Expand Down Expand Up @@ -309,6 +315,7 @@ jobs:
"shuyixiong",
"shyeh25",
"SimengLiu-nv",
"siyidNV",
"sklevtsov-nvidia",
"StanleySun639",
"stnie",
Expand Down Expand Up @@ -410,6 +417,7 @@ jobs:
"zhangcl",
"ZhanruiSunCh",
"zhaoyangwang-nvidia",
"zhaoyuanh-nvidia",
"zhengd-nv",
"zhenhuaw-me",
"zheyuf",
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/bot-command.yml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-FileCopyrightText: Copyright (c) 2024-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
Expand Down Expand Up @@ -52,14 +52,14 @@ jobs:
"`--disable-reuse-test ` *(OPTIONAL)* : Explicitly prevent the pipeline from reusing build artifacts and skipping successful test stages from a previous pipeline. Ensure that all builds and tests are run regardless of previous successes.\n\n" +
"`--disable-fail-fast ` *(OPTIONAL)* : Disable fail fast on build/tests/infra failures.\n\n" +
"`--skip-test ` *(OPTIONAL)* : Skip all test stages, but still run build stages, package stages and sanity check stages. Note: Does **NOT** update GitHub check status.\n\n" +
"`--stage-list \"A10-PyTorch-1, xxx\"` *(OPTIONAL)* : Only run the specified test stages. Supports wildcard `*` for pattern matching (e.g., `\"*PerfSanity*\"` matches all stages containing PerfSanity). Examples: \"A10-PyTorch-1, xxx\", \"*PerfSanity*\". Note: Does **NOT** update GitHub check status.\n\n" +
"`--stage-list \"A10-PyTorch-1, xxx\"` *(OPTIONAL)* : Only run the specified test stages. Supports wildcard `*` for pattern matching (e.g., `\"*PerfSanity*\"` matches all stages containing PerfSanity). Examples: \"A10-PyTorch-1, xxx\", \"*PerfSanity*\". The patterns `\"*\"`, `\"*Post-Merge*\"`, and `\"*PerfSanity*\"`, including equivalent escaped or repeated-star forms and their use in comma-separated lists, require the `ci: post-merge approved` PR label. Note: Does **NOT** update GitHub check status.\n\n" +
"`--gpu-type \"A30, H100_PCIe\"` *(OPTIONAL)* : Only run the test stages on the specified GPU types. Examples: \"A30, H100_PCIe\". Note: Does **NOT** update GitHub check status.\n\n" +
"`--test-backend \"pytorch, cpp\"` *(OPTIONAL)* : Skip test stages which don't match the specified backends. Only support [pytorch, cpp, tensorrt, triton]. Examples: \"pytorch, cpp\" (does not run test stages with tensorrt or triton backend). Note: Does **NOT** update GitHub pipeline status.\n\n" +
"`--only-multi-gpu-test ` *(OPTIONAL)* : Only run the multi-GPU tests. Note: Does **NOT** update GitHub check status.\n\n" +
"`--disable-multi-gpu-test ` *(OPTIONAL)* : Disable the multi-GPU tests. Note: Does **NOT** update GitHub check status.\n\n" +
"`--add-multi-gpu-test ` *(OPTIONAL)* : Force run the multi-GPU tests in addition to running L0 pre-merge pipeline.\n\n" +
"`--post-merge ` *(OPTIONAL)* : Run the L0 post-merge pipeline instead of the ordinary L0 pre-merge pipeline.\n\n" +
"`--extra-stage \"H100_PCIe-TensorRT-Post-Merge-1, xxx\"` *(OPTIONAL)* : Run the ordinary L0 pre-merge pipeline and specified test stages. Supports wildcard `*` for pattern matching. Examples: --extra-stage \"H100_PCIe-TensorRT-Post-Merge-1, xxx\", --extra-stage \"*Post-Merge*\".\n\n" +
"`--post-merge ` *(OPTIONAL)* : Run the L0 post-merge pipeline instead of the ordinary L0 pre-merge pipeline. Requires the `ci: post-merge approved` PR label applied by an active member of `NVIDIA/trt-llm-ci-approvers`. The approval label remains in place when new commits are pushed.\n\n" +
"`--extra-stage \"H100_PCIe-TensorRT-Post-Merge-1, xxx\"` *(OPTIONAL)* : Run the ordinary L0 pre-merge pipeline and specified test stages. Supports wildcard `*` for pattern matching. Examples: --extra-stage \"H100_PCIe-TensorRT-Post-Merge-1, xxx\", --extra-stage \"*Post-Merge*\". The patterns `\"*\"`, `\"*Post-Merge*\"`, and `\"*PerfSanity*\"`, including equivalent escaped or repeated-star forms and their use in comma-separated lists, require the `ci: post-merge approved` PR label.\n\n" +
"`--detailed-log ` *(OPTIONAL)* : Enable flushing out all logs to the Jenkins console. This will significantly increase the log volume and may slow down the job.\n\n" +
"`--debug ` *(OPTIONAL)* : **Experimental feature**. Enable access to the CI container for debugging purpose. Note: Specify exactly one stage in the `stage-list` parameter to access the appropriate container environment. Note: Does **NOT** update GitHub check status.\n\n" +
"`--high-priority ` *(OPTIONAL)* : Run the pipeline with high priority. This option is restricted to authorized users only and will route the job to a high-priority queue.\n\n" +
Expand Down
Loading
Loading