Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
18d47b7
[TRTLLM-16304][feat] In-tree implementation of staircase
WeiHaocheng Sep 10, 2026
8faab78
[TRTLLM-16304][test] Wire the staircase catalog tests into the unit t…
WeiHaocheng Sep 10, 2026
4175d7a
[TRTLLM-16304][test] Add the staircase whole-model gates to the accur…
WeiHaocheng Sep 11, 2026
c9fb814
[TRTLLM-16304][test] Keep the staircase collectives off xdist, drop a…
WeiHaocheng Sep 11, 2026
a566f1b
[TRTLLM-16304][test] Rename the staircase collective rank bodies off …
WeiHaocheng Sep 14, 2026
b52ec33
[TRTLLM-16304][chore] Address the in-tree migration review
WeiHaocheng Sep 15, 2026
16a5e07
[TRTLLM-16304][chore] Rename staircase to modeling_v2
WeiHaocheng Sep 15, 2026
62c771c
[TRTLLM-16304][chore] Follow upstream's removal of MTPDraftModelForCa…
WeiHaocheng Sep 16, 2026
48dec0d
[TRTLLM-16304][fix] Certify thop_attention over two KV pools
WeiHaocheng Sep 16, 2026
2d82d7a
[TRTLLM-16304][infra] Add a CBTS rule for modeling_v2
WeiHaocheng Sep 16, 2026
6ae72b5
[TRTLLM-16304][test] Cut the catalog tests down to what the targets run
WeiHaocheng Sep 16, 2026
f9b71e8
[TRTLLM-16304][chore] Move modeling_v2 under _experimental
WeiHaocheng Sep 17, 2026
d97d768
[TRTLLM-16304][chore] Fix the codespell finding blocking pre-commit
WeiHaocheng Sep 17, 2026
4c9b605
[TRTLLM-16304][test] Carry the modeling_v2 switch on the case, not th…
WeiHaocheng Sep 22, 2026
d43692f
[TRTLLM-16304][test] Move the single-GPU modeling_v2 entries off the …
WeiHaocheng Sep 22, 2026
ff84247
[TRTLLM-16304][test] Put the single-GPU modeling_v2 entries on a stag…
WeiHaocheng Sep 23, 2026
652be0a
[TRTLLM-16304][test] Take the prose back out of the test-db lists
WeiHaocheng Sep 23, 2026
e1c23f4
[TRTLLM-16304][chore] Flatten a target's path into one directory name
WeiHaocheng Sep 23, 2026
8622232
[TRTLLM-16304][chore] Rewrap two lines the flattened path made short …
WeiHaocheng Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -385,6 +385,19 @@
/tensorrt_llm/scaffolding @WeiHaocheng @dc3671
/tests/unittest/scaffolding @WeiHaocheng @dc3671

# ===== MODELING_V2 =====
# Overrides the /tensorrt_llm/_torch runtime-devs rule above for this subtree.
# Individual handles rather than a team, like SCAFFOLDING: this is one bounded
# experiment with named owners, not a standing domain.
/tensorrt_llm/_torch/_experimental/modeling_v2 @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju
/tests/unittest/_torch/modeling_v2 @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju
/tests/integration/defs/accuracy/test_modeling_v2_deepseek_v3.py @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju
/tests/integration/defs/accuracy/test_modeling_v2_gpt_oss.py @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju
# The two catalog categories whose contracts state kernel behaviour the
# attention and MoE owners are the authority on; co-owned rather than reassigned.
/tensorrt_llm/_torch/_experimental/modeling_v2/catalog/attention @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju @xxi-nv @yuxianq
/tensorrt_llm/_torch/_experimental/modeling_v2/catalog/moe @tianyuxbear @Wanli-Jiang @WeiHaocheng @litaotju @xxi-nv @yuxianq

## TensorRT-LLM LLM Disaggregated
/examples/disaggregated @NVIDIA/trt-llm-disagg-devs @NVIDIA/trt-llm-doc-owners
/examples/disaggregated/slurm/benchmark @NVIDIA/trt-llm-disagg-devs @NVIDIA/trtllm-bench-reviewers
Expand Down
6 changes: 4 additions & 2 deletions jenkins/scripts/cbts/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ filter chain.

## Rules

Nine rules, registered in `main.py::RULE_CLASSES`:
Ten rules, registered in `main.py::RULE_CLASSES`:

| Rule | Scope | Files |
|---|---|---|
Expand All @@ -46,6 +46,7 @@ Nine rules, registered in `main.py::RULE_CLASSES`:
| `TestListRule` | `testlistonly` | `tests/integration/test_lists/test-db/*.yml` |
| `VisualGenRule` | `visualgenonly` | `examples/visual_gen/**`, `scripts/visualgen_eval/**`, `tensorrt_llm/_torch/visual_gen/**`, `tensorrt_llm/media/**`, `tensorrt_llm/visual_gen/**` (excl. `.md`; reference images such as `cat_piano.png` ARE test fixtures and stay claimed; outward-facing files force fallback) |
| `SpecDecRule` | `specdeconly` | `tensorrt_llm/_torch/speculative/**`, `tensorrt_llm/models/{eagle,medusa,redrafter}/**`, `examples/{eagle,medusa,redrafter,draft_target_model,ngram}/**`, `examples/llm-api/llm_speculative_decoding.py` (excl. `.md`; other suffixes incl. images kept as potential test fixtures) |
| `ModelingV2Rule` | `modelingv2only` | `tensorrt_llm/_torch/_experimental/modeling_v2/**` (excl. `.md`) |
| `AgentFlowRule` | `agentflowonly` | `agent-flow/**` (excl. `.md`) |
| `OpenEngineRule` | `openengineonly` | `tensorrt_llm/grpc/openengine/**` (excl. `.md`) |
| `OutOfScopeRule` | `noop` | QA / dev test lists, `.test_durations`, `microbenchmarks/`, `**/*.md` (image suffixes intentionally not claimed — image fixtures cannot be distinguished from doc diagrams by location, so image edits fall back to baseline) |
Expand All @@ -61,9 +62,10 @@ See `rules/README.md` for per-rule logic.
| `testlistonly` | `TestListRule` fired solo: PR only adds entries under `tests/integration/test_lists/test-db/*.yml`. |
| `visualgenonly` | `VisualGenRule` fired solo: PR only touches VisualGen internal source paths (`examples/visual_gen/**`, `scripts/visualgen_eval/**`, `tensorrt_llm/_torch/visual_gen/**`; excl. `.md`; image fixtures like `cat_piano.png` are claimed). Narrows to blocks containing VG test entries. Outward-facing files under `tensorrt_llm/visual_gen/**` and `tensorrt_llm/media/**` (eagerly imported by `trtllm-serve`) force `null` fallback. |
| `specdeconly` | `SpecDecRule` fired solo: PR only touches speculative-decoding source paths (`tensorrt_llm/_torch/speculative/**`, `tensorrt_llm/models/{eagle,medusa,redrafter}/**`, `examples/{eagle,medusa,redrafter,draft_target_model,ngram}/**`, `examples/llm-api/llm_speculative_decoding.py`; excl. `.md`). Narrows to blocks containing spec-dec test entries (eagle / medusa / redrafter / ngram / draft-target-model / MTP). |
| `modelingv2only` | `ModelingV2Rule` fired solo: PR only touches the modeling_v2 subtree (`tensorrt_llm/_torch/_experimental/modeling_v2/**`; excl. `.md`, which is a fifth of the subtree — every catalog entry carries a contract document). Narrows to blocks containing modeling_v2 test entries (`unittest/_torch/modeling_v2/`, `test_modeling_v2_*`). No outward-facing fallback is needed: nothing imports the subtree unless `TRTLLM_MODELING_V2` is set, and its one caller outside the subtree (`_torch/models/modeling_auto.py`) is left unclaimed, so touching the shared resolver falls back to baseline. |
| `agentflowonly` | `AgentFlowRule` fired solo: PR only touches `agent-flow/**` source or test files (excl. `.md`). Runs `CPU-AgentFlow-UnitTest`. |
| `openengineonly` | `OpenEngineRule` fired solo: PR only touches `tensorrt_llm/grpc/openengine/**` source files (excl. `.md`). Narrows to the registered OpenEngine unit tests: the stub-based ones on the always-run `CPU-Generic-*` stages, plus `test_capability_conformance.py` on `A10-PyTorch-*`, which needs a GPU. |
| `testsonly` | Multiple rules from the testsonly family fired (`waiveonly`, `testdefonly`, `testlistonly`, `visualgenonly`, `specdeconly`, `agentflowonly`, `openengineonly`); their narrows union. |
| `testsonly` | Multiple rules from the testsonly family fired (`waiveonly`, `testdefonly`, `testlistonly`, `visualgenonly`, `specdeconly`, `modelingv2only`, `agentflowonly`, `openengineonly`); their narrows union. |
| `noop` | Rule(s) fired but determined no test stages need to run (QA-only path, removals-only test list, all-miss waives, in-namespace .py with no covering YAML entry, docs-only edits). Layer 2 still applies. |
| `null` (fallback) | A rule cannot decide, scopes don't combine, or there are unhandled files. Groovy defers to baseline filter chain. |

Expand Down
4 changes: 4 additions & 0 deletions jenkins/scripts/cbts/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@
from rules._helpers import strip_noop_diff_lines # noqa: E402
from rules.agent_flow_rule import AgentFlowRule # noqa: E402
from rules.base import PRInputs, Rule, RuleResult, format_reason # noqa: E402
from rules.modeling_v2_rule import ModelingV2Rule # noqa: E402
from rules.openengine_rule import OpenEngineRule # noqa: E402
from rules.out_of_scope_rule import OutOfScopeRule # noqa: E402
from rules.spec_dec_rule import SpecDecRule # noqa: E402
Expand All @@ -81,6 +82,7 @@
TestListRule,
VisualGenRule,
SpecDecRule,
ModelingV2Rule,
AgentFlowRule,
OpenEngineRule,
OutOfScopeRule,
Expand All @@ -98,6 +100,7 @@ def build_rules(
TestListRule(yaml_index, stages, repo_root=repo_root),
VisualGenRule(yaml_index, stages),
SpecDecRule(yaml_index, stages),
ModelingV2Rule(yaml_index, stages),
AgentFlowRule(yaml_index, stages),
OpenEngineRule(yaml_index, stages),
OutOfScopeRule(yaml_index, stages),
Expand Down Expand Up @@ -217,6 +220,7 @@ def _rule_reason(rule, r) -> dict:
"testlistonly",
"visualgenonly",
"specdeconly",
"modelingv2only",
"agentflowonly",
"openengineonly",
}
Expand Down
51 changes: 51 additions & 0 deletions jenkins/scripts/cbts/rules/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ for the overall CBTS architecture.
| `test_list_rule.py` | `TestListRule` | `testlistonly` | `tests/integration/test_lists/test-db/*.yml` |
| `visual_gen_rule.py` | `VisualGenRule` | `visualgenonly` | `examples/visual_gen/**`, `scripts/visualgen_eval/**`, `tensorrt_llm/_torch/visual_gen/**`, `tensorrt_llm/media/**`, `tensorrt_llm/visual_gen/**` (each excl. `.md`) |
| `spec_dec_rule.py` | `SpecDecRule` | `specdeconly` | `tensorrt_llm/_torch/speculative/**`, `tensorrt_llm/models/{eagle,medusa,redrafter}/**`, `examples/{eagle,medusa,redrafter,draft_target_model,ngram}/**`, `examples/llm-api/llm_speculative_decoding.py` (each excl. `.md`) |
| `modeling_v2_rule.py` | `ModelingV2Rule` | `modelingv2only` | `tensorrt_llm/_torch/_experimental/modeling_v2/**` (excl. `.md`) |
| `agent_flow_rule.py` | `AgentFlowRule` | `agentflowonly` | `agent-flow/**` (excl. `.md`) → the single `CPU-AgentFlow-UnitTest` stage; not test-db-driven |
| `openengine_rule.py` | `OpenEngineRule` | `openengineonly` | `tensorrt_llm/grpc/openengine/**` (excl. `.md`) → the `l0_cpu` block containing `unittest/grpc/openengine/` |
| `out_of_scope_rule.py` | `OutOfScopeRule` | `noop` | `tests/integration/test_lists/{qa,dev}/**`, `tests/integration/defs/.test_durations*`, `tests/microbenchmarks/**`, `**/*.md` (image suffixes intentionally not claimed — fall back to baseline since fixtures and doc diagrams are indistinguishable by location) |
Expand Down Expand Up @@ -254,6 +255,56 @@ Outcomes:
- Spec-dec source touched but no spec-dec block found anywhere
(defensive) → `scope=None` (fallback).

## ModelingV2Rule

Path-only rule. Claims non-documentation source changes under
`tensorrt_llm/_torch/_experimental/modeling_v2/`, the self-contained second modeling
path (one flat forward per checkpoint/arch/parallel triple, assembled
from a catalog of op wrappers).

`.md` exclusion carries more weight here than elsewhere: every catalog
entry ships a contract document, so roughly a fifth of the subtree is
Markdown. Claiming those would let a documentation-only PR pull in
multi-GPU GB300 stages. Other suffixes are NOT excluded — a data file
under the subtree could be a fixture, so the rule keeps claiming it
(safe over-run).

Block selection — entry-pattern based only:
modeling_v2 has no `condition.terms.backend` of its own; its entries sit
in `backend: pytorch` blocks beside everything else. A block belongs to
modeling_v2 iff one of its `tests:` entries matches
`_MV2_ENTRY_PATTERNS`:

- `unittest/_torch/modeling_v2/` — the op-level catalog matrix, carried
as whole-directory entries (one on `l0_gb300`, one on
`l0_gb300_multi_gpus` for the 4-rank collectives).
- `test_modeling_v2_` — the accuracy gates, and any future unit file.

Both markers are exact by construction rather than by luck: every test
file in the subtree is named `test_modeling_v2_*` precisely so it cannot
collide with the upstream test of the same op. So unlike `SpecDecRule`'s
`mtp_nextn`, no substring here can claim an unrelated entry and no
carve-out is needed.

Outward fallback: not needed, and by design rather than by accident.
Nothing imports the subtree unless `TRTLLM_MODELING_V2` is set —
`AutoModelForCausalLM._resolve_class` calls `modeling_v2_resolve`, which
returns immediately when the switch is off, and the routing modules are
imported lazily behind it. The one caller outside the subtree,
`tensorrt_llm/_torch/models/modeling_auto.py`, is deliberately left
unclaimed: a change to the shared resolver falls back to baseline, which
is what it deserves.

`sanity_relevant=False` — the subtree ships no user-facing entry point
and is not imported by `trtllm-serve` or by `import tensorrt_llm`, so
none of it is what PackageSanityCheck verifies about the wheel.
`perfsanity_relevant` is dynamic (True only if a matched block lives in
a `*_perf_sanity*` yaml); there are no modeling_v2 perf-sanity entries
today, so it aggregates to False.

Source changed but no modeling_v2 block in any yaml (defensive) →
`scope=None` (fallback).

## OpenEngineRule

Claims non-documentation source changes under `tensorrt_llm/grpc/openengine/` and keeps only test-db
Expand Down
166 changes: 166 additions & 0 deletions jenkins/scripts/cbts/rules/modeling_v2_rule.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
Comment thread
WeiHaocheng marked this conversation as resolved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""ModelingV2Rule — narrows CI when the modeling_v2 subtree changes.

modeling_v2 is a second modeling path living entirely under
`tensorrt_llm/_torch/_experimental/modeling_v2/`: one self-contained forward per
(checkpoint, GPU arch, parallel topology), assembled from a catalog of
op wrappers.

Block selection — entry-pattern based only:
It has no `condition.terms.backend` of its own; its entries sit in
`backend: pytorch` blocks beside everything else. A block belongs to
modeling_v2 iff one of its `tests:` entries matches a marker in
`_MV2_ENTRY_PATTERNS`. Those markers are exact by construction rather
than by luck: every unit test file in the subtree is named
`test_modeling_v2_*` precisely so it cannot collide with the upstream
test of the same op, and the accuracy files follow the same prefix. So
there is no substring that could claim an unrelated entry, and no
`mtp_nextn=0`-style carve-out is needed.

Outward fallback: not needed, and that is a property of the design
rather than an accident. Nothing imports this subtree unless
`TRTLLM_MODELING_V2` is set: `AutoModelForCausalLM._resolve_class` calls
`modeling_v2_resolve`, which returns immediately when the switch is off,
and the routing modules are imported lazily behind it. The one caller
outside the subtree is `tensorrt_llm/_torch/models/modeling_auto.py`,
which this rule does not claim -- a PR touching it falls back to
baseline, which is what a change to the shared resolver deserves.

`.md` exclusion matters more here than for most rules: the catalog
carries a contract document per entry, so roughly a fifth of the files
in the subtree are Markdown. Claiming them would make a
documentation-only PR pull in multi-GPU GB300 stages.

PerfSanity policy: `perfsanity_relevant` is dynamic, True only when a
matched block lives in a `*_perf_sanity*` yaml -- same as AutoDeployRule
/ VisualGenRule / SpecDecRule. modeling_v2 has no perf-sanity entries
today, so this aggregates to False and Groovy Layer 2 drops the
force-keep of `*-PerfSanity-*` stages.

Sanity policy: `sanity_relevant=False`. The subtree ships no
user-facing entry point and is not imported by `trtllm-serve` or by
`import tensorrt_llm`, so nothing it contains is what PackageSanityCheck
verifies about the wheel.
"""

from __future__ import annotations

from typing import Optional

from blocks import Stage, YAMLIndex, _entry_target

from ._helpers import resolve_affected_stages, stages_by_yaml_stem
from .base import PRInputs, Rule, RuleResult

# Source-path prefixes the rule may claim. Tests under tests/** are left
# to TestsDefRule; the two scopes combine via _TESTSONLY_FAMILY.
_MV2_SRC_PREFIXES: tuple[str, ...] = ("tensorrt_llm/_torch/_experimental/modeling_v2/",)

# Substrings that mark a test entry as modeling_v2. Both are unambiguous:
# - "unittest/_torch/modeling_v2/" → the op-level catalog matrix, taken
# as whole-directory entries (one on l0_gb300, one on
# l0_gb300_multi_gpus for the 4-rank collectives)
# - "test_modeling_v2_" → the accuracy gates, and any future unit file
# named by the subtree's own convention
_MV2_ENTRY_PATTERNS: tuple[str, ...] = (
"unittest/_torch/modeling_v2/",
"test_modeling_v2_",
)


def _is_mv2_claim(path: str) -> bool:
"""Decide whether ModelingV2Rule claims `path`.

`*.md` is excluded so a contract-only edit does not force GPU stages
-- `OutOfScopeRule` claims those as noop instead. Other suffixes are
NOT excluded: a data file under this subtree could be a fixture, so
the rule keeps claiming it and re-runs the stages (safe over-run).
"""
if not path.startswith(_MV2_SRC_PREFIXES):
return False
if path.endswith(".md"):
return False
return True


def _entry_is_mv2(entry: str) -> bool:
return any(p in entry for p in _MV2_ENTRY_PATTERNS)


def _mv2_entries(block) -> list[str]:
return [t for t in block.tests if _entry_is_mv2(t)]


def _is_perf_sanity_stem(stem: str) -> bool:
"""True for perf-sanity yaml stems (`l0_*_perf_sanity*`)."""
return "perf_sanity" in stem


class ModelingV2Rule(Rule):
name = "modelingv2"
needs_diff_for: tuple[str, ...] = ()

def __init__(self, yaml_index: YAMLIndex, stages: dict[str, Stage]) -> None:
self.yaml_index = yaml_index
self._stages_by_yaml = stages_by_yaml_stem(stages)

def apply(self, pr: PRInputs) -> Optional[RuleResult]:
claimed = {f for f in pr.changed_files if _is_mv2_claim(f)}
if not claimed:
return None

block_filters: dict[tuple[str, int], dict[str, set[str]]] = {}
for block in self.yaml_index.blocks:
entries = _mv2_entries(block)
if not entries:
continue
key = (block.yaml_stem, block.block_index)
prefix_dict = block_filters.setdefault(key, {})
for entry in entries:
target = _entry_target(entry)
if target:
prefix_dict.setdefault(target, set()).add(entry)

if not block_filters:
# Defensive: modeling_v2 source changed but no modeling_v2 block
# exists in any yaml. Do not fabricate stages -- fall back to
# baseline so the change still gets coverage. Reachable if the
# subtree's entries are ever removed from the test-db without
# the subtree going with them.
return RuleResult(
handled_files=claimed,
affected_stages=set(),
scope=None,
reason=(
f"modelingv2: {len(claimed)} modeling_v2 source file(s); "
"no modeling_v2 block matched in any test-db yaml — fallback"
),
)

affected = resolve_affected_stages(block_filters, self.yaml_index, self._stages_by_yaml)
perfsanity_relevant = any(_is_perf_sanity_stem(stem) for stem, _ in block_filters)

return RuleResult(
handled_files=claimed,
affected_stages=affected,
scope="modelingv2only",
block_filters=block_filters,
sanity_relevant=False,
perfsanity_relevant=perfsanity_relevant,
reason=(
f"modelingv2: {len(claimed)} modeling_v2 source file(s) → "
f"{len(block_filters)} modeling_v2 block(s), {len(affected)} stage(s)"
),
)
15 changes: 15 additions & 0 deletions tensorrt_llm/_torch/_experimental/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Subpackages here are not covered by the API stability tests.

Anything under this package may change shape or be removed without a
deprecation cycle. Import it from outside `tensorrt_llm` at your own risk;
in-tree callers should reach it through a switch that stays off by default,
the way `modeling_v2` is reached through `TRTLLM_MODELING_V2`.

Nothing is re-exported here on purpose, so this module itself pulls in no
subpackage. That is not the same as saying a subtree here is unreachable at
startup -- `modeling_v2` is imported eagerly by `_torch/models/modeling_auto.py`
-- only that reaching one has to be written down at the import site rather than
happening as a side effect of this package.
"""
Loading
Loading