-
Notifications
You must be signed in to change notification settings - Fork 2.8k
[TRTLLM-16304][feat] In-tree implementation of staircase #19056
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
WeiHaocheng
merged 19 commits into
NVIDIA:main
from
WeiHaocheng:feat/staircase-upstream
Sep 27, 2026
Merged
Changes from all commits
Commits
Show all changes
19 commits
Select commit
Hold shift + click to select a range
18d47b7
[TRTLLM-16304][feat] In-tree implementation of staircase
WeiHaocheng 8faab78
[TRTLLM-16304][test] Wire the staircase catalog tests into the unit t…
WeiHaocheng 4175d7a
[TRTLLM-16304][test] Add the staircase whole-model gates to the accur…
WeiHaocheng c9fb814
[TRTLLM-16304][test] Keep the staircase collectives off xdist, drop a…
WeiHaocheng a566f1b
[TRTLLM-16304][test] Rename the staircase collective rank bodies off …
WeiHaocheng b52ec33
[TRTLLM-16304][chore] Address the in-tree migration review
WeiHaocheng 16a5e07
[TRTLLM-16304][chore] Rename staircase to modeling_v2
WeiHaocheng 62c771c
[TRTLLM-16304][chore] Follow upstream's removal of MTPDraftModelForCa…
WeiHaocheng 48dec0d
[TRTLLM-16304][fix] Certify thop_attention over two KV pools
WeiHaocheng 2d82d7a
[TRTLLM-16304][infra] Add a CBTS rule for modeling_v2
WeiHaocheng 6ae72b5
[TRTLLM-16304][test] Cut the catalog tests down to what the targets run
WeiHaocheng f9b71e8
[TRTLLM-16304][chore] Move modeling_v2 under _experimental
WeiHaocheng d97d768
[TRTLLM-16304][chore] Fix the codespell finding blocking pre-commit
WeiHaocheng 4c9b605
[TRTLLM-16304][test] Carry the modeling_v2 switch on the case, not th…
WeiHaocheng d43692f
[TRTLLM-16304][test] Move the single-GPU modeling_v2 entries off the …
WeiHaocheng ff84247
[TRTLLM-16304][test] Put the single-GPU modeling_v2 entries on a stag…
WeiHaocheng 652be0a
[TRTLLM-16304][test] Take the prose back out of the test-db lists
WeiHaocheng e1c23f4
[TRTLLM-16304][chore] Flatten a target's path into one directory name
WeiHaocheng 8622232
[TRTLLM-16304][chore] Rewrap two lines the flattened path made short …
WeiHaocheng File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,166 @@ | ||
| # SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | ||
| # | ||
| # Licensed under the Apache License, Version 2.0 (the "License"); | ||
| # you may not use this file except in compliance with the License. | ||
| # You may obtain a copy of the License at | ||
| # | ||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||
| # | ||
| # Unless required by applicable law or agreed to in writing, software | ||
| # distributed under the License is distributed on an "AS IS" BASIS, | ||
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
| """ModelingV2Rule — narrows CI when the modeling_v2 subtree changes. | ||
|
|
||
| modeling_v2 is a second modeling path living entirely under | ||
| `tensorrt_llm/_torch/_experimental/modeling_v2/`: one self-contained forward per | ||
| (checkpoint, GPU arch, parallel topology), assembled from a catalog of | ||
| op wrappers. | ||
|
|
||
| Block selection — entry-pattern based only: | ||
| It has no `condition.terms.backend` of its own; its entries sit in | ||
| `backend: pytorch` blocks beside everything else. A block belongs to | ||
| modeling_v2 iff one of its `tests:` entries matches a marker in | ||
| `_MV2_ENTRY_PATTERNS`. Those markers are exact by construction rather | ||
| than by luck: every unit test file in the subtree is named | ||
| `test_modeling_v2_*` precisely so it cannot collide with the upstream | ||
| test of the same op, and the accuracy files follow the same prefix. So | ||
| there is no substring that could claim an unrelated entry, and no | ||
| `mtp_nextn=0`-style carve-out is needed. | ||
|
|
||
| Outward fallback: not needed, and that is a property of the design | ||
| rather than an accident. Nothing imports this subtree unless | ||
| `TRTLLM_MODELING_V2` is set: `AutoModelForCausalLM._resolve_class` calls | ||
| `modeling_v2_resolve`, which returns immediately when the switch is off, | ||
| and the routing modules are imported lazily behind it. The one caller | ||
| outside the subtree is `tensorrt_llm/_torch/models/modeling_auto.py`, | ||
| which this rule does not claim -- a PR touching it falls back to | ||
| baseline, which is what a change to the shared resolver deserves. | ||
|
|
||
| `.md` exclusion matters more here than for most rules: the catalog | ||
| carries a contract document per entry, so roughly a fifth of the files | ||
| in the subtree are Markdown. Claiming them would make a | ||
| documentation-only PR pull in multi-GPU GB300 stages. | ||
|
|
||
| PerfSanity policy: `perfsanity_relevant` is dynamic, True only when a | ||
| matched block lives in a `*_perf_sanity*` yaml -- same as AutoDeployRule | ||
| / VisualGenRule / SpecDecRule. modeling_v2 has no perf-sanity entries | ||
| today, so this aggregates to False and Groovy Layer 2 drops the | ||
| force-keep of `*-PerfSanity-*` stages. | ||
|
|
||
| Sanity policy: `sanity_relevant=False`. The subtree ships no | ||
| user-facing entry point and is not imported by `trtllm-serve` or by | ||
| `import tensorrt_llm`, so nothing it contains is what PackageSanityCheck | ||
| verifies about the wheel. | ||
| """ | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| from typing import Optional | ||
|
|
||
| from blocks import Stage, YAMLIndex, _entry_target | ||
|
|
||
| from ._helpers import resolve_affected_stages, stages_by_yaml_stem | ||
| from .base import PRInputs, Rule, RuleResult | ||
|
|
||
| # Source-path prefixes the rule may claim. Tests under tests/** are left | ||
| # to TestsDefRule; the two scopes combine via _TESTSONLY_FAMILY. | ||
| _MV2_SRC_PREFIXES: tuple[str, ...] = ("tensorrt_llm/_torch/_experimental/modeling_v2/",) | ||
|
|
||
| # Substrings that mark a test entry as modeling_v2. Both are unambiguous: | ||
| # - "unittest/_torch/modeling_v2/" → the op-level catalog matrix, taken | ||
| # as whole-directory entries (one on l0_gb300, one on | ||
| # l0_gb300_multi_gpus for the 4-rank collectives) | ||
| # - "test_modeling_v2_" → the accuracy gates, and any future unit file | ||
| # named by the subtree's own convention | ||
| _MV2_ENTRY_PATTERNS: tuple[str, ...] = ( | ||
| "unittest/_torch/modeling_v2/", | ||
| "test_modeling_v2_", | ||
| ) | ||
|
|
||
|
|
||
| def _is_mv2_claim(path: str) -> bool: | ||
| """Decide whether ModelingV2Rule claims `path`. | ||
|
|
||
| `*.md` is excluded so a contract-only edit does not force GPU stages | ||
| -- `OutOfScopeRule` claims those as noop instead. Other suffixes are | ||
| NOT excluded: a data file under this subtree could be a fixture, so | ||
| the rule keeps claiming it and re-runs the stages (safe over-run). | ||
| """ | ||
| if not path.startswith(_MV2_SRC_PREFIXES): | ||
| return False | ||
| if path.endswith(".md"): | ||
| return False | ||
| return True | ||
|
|
||
|
|
||
| def _entry_is_mv2(entry: str) -> bool: | ||
| return any(p in entry for p in _MV2_ENTRY_PATTERNS) | ||
|
|
||
|
|
||
| def _mv2_entries(block) -> list[str]: | ||
| return [t for t in block.tests if _entry_is_mv2(t)] | ||
|
|
||
|
|
||
| def _is_perf_sanity_stem(stem: str) -> bool: | ||
| """True for perf-sanity yaml stems (`l0_*_perf_sanity*`).""" | ||
| return "perf_sanity" in stem | ||
|
|
||
|
|
||
| class ModelingV2Rule(Rule): | ||
| name = "modelingv2" | ||
| needs_diff_for: tuple[str, ...] = () | ||
|
|
||
| def __init__(self, yaml_index: YAMLIndex, stages: dict[str, Stage]) -> None: | ||
| self.yaml_index = yaml_index | ||
| self._stages_by_yaml = stages_by_yaml_stem(stages) | ||
|
|
||
| def apply(self, pr: PRInputs) -> Optional[RuleResult]: | ||
| claimed = {f for f in pr.changed_files if _is_mv2_claim(f)} | ||
| if not claimed: | ||
| return None | ||
|
|
||
| block_filters: dict[tuple[str, int], dict[str, set[str]]] = {} | ||
| for block in self.yaml_index.blocks: | ||
| entries = _mv2_entries(block) | ||
| if not entries: | ||
| continue | ||
| key = (block.yaml_stem, block.block_index) | ||
| prefix_dict = block_filters.setdefault(key, {}) | ||
| for entry in entries: | ||
| target = _entry_target(entry) | ||
| if target: | ||
| prefix_dict.setdefault(target, set()).add(entry) | ||
|
|
||
| if not block_filters: | ||
| # Defensive: modeling_v2 source changed but no modeling_v2 block | ||
| # exists in any yaml. Do not fabricate stages -- fall back to | ||
| # baseline so the change still gets coverage. Reachable if the | ||
| # subtree's entries are ever removed from the test-db without | ||
| # the subtree going with them. | ||
| return RuleResult( | ||
| handled_files=claimed, | ||
| affected_stages=set(), | ||
| scope=None, | ||
| reason=( | ||
| f"modelingv2: {len(claimed)} modeling_v2 source file(s); " | ||
| "no modeling_v2 block matched in any test-db yaml — fallback" | ||
| ), | ||
| ) | ||
|
|
||
| affected = resolve_affected_stages(block_filters, self.yaml_index, self._stages_by_yaml) | ||
| perfsanity_relevant = any(_is_perf_sanity_stem(stem) for stem, _ in block_filters) | ||
|
|
||
| return RuleResult( | ||
| handled_files=claimed, | ||
| affected_stages=affected, | ||
| scope="modelingv2only", | ||
| block_filters=block_filters, | ||
| sanity_relevant=False, | ||
| perfsanity_relevant=perfsanity_relevant, | ||
| reason=( | ||
| f"modelingv2: {len(claimed)} modeling_v2 source file(s) → " | ||
| f"{len(block_filters)} modeling_v2 block(s), {len(affected)} stage(s)" | ||
| ), | ||
| ) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,15 @@ | ||
| # SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | ||
| # SPDX-License-Identifier: Apache-2.0 | ||
| """Subpackages here are not covered by the API stability tests. | ||
|
|
||
| Anything under this package may change shape or be removed without a | ||
| deprecation cycle. Import it from outside `tensorrt_llm` at your own risk; | ||
| in-tree callers should reach it through a switch that stays off by default, | ||
| the way `modeling_v2` is reached through `TRTLLM_MODELING_V2`. | ||
|
|
||
| Nothing is re-exported here on purpose, so this module itself pulls in no | ||
| subpackage. That is not the same as saying a subtree here is unreachable at | ||
| startup -- `modeling_v2` is imported eagerly by `_torch/models/modeling_auto.py` | ||
| -- only that reaching one has to be written down at the import site rather than | ||
| happening as a side effect of this package. | ||
| """ |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.