Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/TensorRT-LLM/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughThe FMHA CMake configuration now prunes architectures 100, 103, and 100f. It retains architecture 107 in CUDA source filtering and cubin archive selection, and removes the directory-level ChangesFMHA Architecture Pruning
Estimated code review effort: 2 (Simple) | ~8 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to The change preserves the SM107 context FMHA cubins needed for Rubin while retaining the other listed pruning. No actionable merge risk is evident in the reviewed scope. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
aa7c162 to
a709ddf
Compare
|
Automatically added "ci: full pre-merge approved" because this PR has satisfied the required GitHub review approvals. Unresolved review conversations and other required checks remain independent merge requirements. |
|
/bot run |
|
PR_Github #75617 [ run ] triggered by Bot. Commit: |
|
PR_Github #75617 [ run ] completed with state
|
|
Evaluate behavioral incompatibilities involving the supplied PR changes and the Verify the three full commit IDs and their merge-base. For divergent branches, When head contains target ( First identify changed or removed interfaces and symbols, then search repository Classify findings in prose as cross-branch interactions or PR-local compatibility Use read-only source and Git inspection. Do not modify checked-out files, execute Assess independent findings separately. Uncertainty about one path does not
Write exactly one standalone heading
Then give the verdict, findings or inspected paths, and limitations. For PASS and Copy the supplied identity and commit IDs verbatim. Always include the result Analyze only these fixed revisions in NVIDIA/TensorRT-LLM, not the hosting PR: |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -u
H=a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25
T=af131fc340c62e77f5e5238f31c9bb7cb4364307
B=b91158667191d2226006e6b20557bcba2d02d2a1
for x in "$H" "$T" "$B"; do echo "COMMIT $x"; git cat-file -t "$x" || :; git rev-parse "$x^{commit}" || :; done
echo BASE; git merge-base "$H" "$T" || :
echo HEAD_DIFF; git diff --stat "$B" "$H"; git diff --name-status "$B" "$H"
echo TARGET_DIFF; git diff --stat "$B" "$T"; git diff --name-status "$B" "$T"
echo HEAD_CHANGE; git diff "$B" "$H" -- cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/CMakeLists.txt
echo TARGET_CHANGE; git diff "$B" "$T" -- cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/CMakeLists.txtLength of output: 201352 🏁 Script executed: #!/bin/bash
set -u
H=a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25
T=af131fc340c62e77f5e5238f31c9bb7cb4364307
B=b91158667191d2226006e6b20557bcba2d02d2a1
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo MERGE_PREVIEW
git merge-tree "$B" "$H" "$T" | head -110
echo RELATED_CHANGED_NAMES
git diff --name-only "$B" "$T" -- cpp/tensorrt_llm | grep -vE '(/cubin/|/lib/)' | grep -Ei '(fmha|attention|cubin|CMake)' | head -100
echo TARGET_META_SUMMARY
git diff --numstat "$B" "$T" -- "$P/kernelMetaInfo.h"
git diff "$B" "$T" -- "$P/kernelMetaInfo.h" | head -130
echo HEAD_CMAKE
git show "$H:$P/CMakeLists.txt" | nl -ba | head -125
echo TARGET_FMHA_RELATED
git diff "$B" "$T" -- cpp/tensorrt_llm/kernels/contextFusedMultiHeadAttention/fmhaRunner.cpp "$P/trtllmGen_fmha_export/FmhaAutoTuner.h" "$P/trtllmGen_fmha_export/FmhaOptions.h" | head -260
echo ARCHIVE_PATHS
fd -i 'tllm_cubin_archive|filter_source_cuda_architectures' . | head -30Length of output: 22584 🏁 Script executed: #!/bin/bash
set -u
H=a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25
T=af131fc340c62e77f5e5238f31c9bb7cb4364307
B=b91158667191d2226006e6b20557bcba2d02d2a1
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo THREE_WAY_RELEVANT
git merge-tree "$B" "$H" "$T" | grep -n -C4 -E 'fmha/CMakeLists.txt|kernelMetaInfo.h|CONFLICT' | head -100
echo TARGET_CMAKE_SHA
for c in "$B" "$H" "$T"; do git rev-parse "$c:$P/CMakeLists.txt"; done
echo TARGET_ARCHIVE_META
git diff --raw "$B" "$T" -- "$P/kernelMetaInfo.h" | head
git ls-tree "$T" "$P/kernelMetaInfo.h"
git show "$T:$P/kernelMetaInfo.h" | head -50
echo HELPER_FILE
ast-grep outline cpp/cmake/modules/tllm_cubin_archive.cmake --items all | head -100
nl -ba cpp/cmake/modules/tllm_cubin_archive.cmake | head -260
echo REFERENCES_WORKTREE
rg -n -C2 'EXCLUDE_SM_107|filter_source_cuda_architectures|tllm_add_cubin_archive_sources' cpp/cmake cpp/tensorrt_llm -g '*.cmake' -g 'CMakeLists.txt' | head -200
echo META_REFERENCES
rg -n -C2 'kernelMetaInfo\.h|EXCLUDE_SM_107|Sm107a|sm107a' "$P" -g '!*.cubin.tar.zst' | head -150Length of output: 51320 🏁 Script executed: #!/bin/bash
set -u
B=b91158667191d2226006e6b20557bcba2d02d2a1
H=a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25
T=af131fc340c62e77f5e5238f31c9bb7cb4364307
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo TARGET_META_DIFF
git diff "$B" "$T" -- "$P/cubin/kernelMetaInfo.h" | head -145
echo FILTER
sed -n '590,690p' cpp/cmake/modules/cuda_configuration.cmake | nl -ba -v590
echo PARENT_FLAGS
rg -n -C3 'EXCLUDE_SM_107|add_subdirectory\(.*fmha|add_compile_definitions|remove_definitions' cpp/CMakeLists.txt cpp/tensorrt_llm/kernels/CMakeLists.txt "$P/CMakeLists.txt" cpp/cmake/modules/cuda_configuration.cmake | head -120
echo SYMBOL_ANALYSIS
python - <<'PY'
import subprocess,re
p='cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/'
for c in ('b91158667191d2226006e6b20557bcba2d02d2a1','a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25','af131fc340c62e77f5e5238f31c9bb7cb4364307'):
text=subprocess.check_output(['git','show',f'{c}:{p}kernelMetaInfo.h'],text=True)
archives=subprocess.check_output(['git','ls-tree','-r','--name-only',c,p],text=True).splitlines()
names={x.rsplit('/',1)[-1][:-len('.cubin.tar.zst')] for x in archives if x.endswith('.cubin.tar.zst')}
symbols={x for x in re.findall(r'\b(FmhaSm107aKernel_\w+)_cubin\b',text)}
missing=sorted(symbols-names)
print(c,'archive count',len(names),'sm107 archives',sum(x.startswith('FmhaSm107aKernel_') for x in names),'metadata symbols',len(symbols),'not in archive',len(missing),missing[:8])
# inspect references beyond extern:
print('metadata with initializer?',len(re.findall(r'FmhaSm107aKernel_\w+_cubin\s*,',text)))
PYLength of output: 30455 🏁 Script executed: #!/bin/bash
set -u
B=b91158667191d2226006e6b20557bcba2d02d2a1
T=af131fc340c62e77f5e5238f31c9bb7cb4364307
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo FILTER_END
sed -n '687,710p' cpp/cmake/modules/cuda_configuration.cmake | nl -ba -v687
echo GATES_FULL_SCOPE
rg -n -C2 'EXCLUDE_SM_107|add_definitions\(' cpp/cmake cpp/tensorrt_llm/kernels/trtllmGenKernels cpp/CMakeLists.txt -g '*.cmake' -g 'CMakeLists.txt' | grep -A4 -B3 -E 'EXCLUDE_SM_107|add_definitions' | head -120
echo FMHA_RUNTIME
rg -n -C4 'getSMVersion|mSM|Sm107a|mKernelMeta|sKernelMeta|kernelMetaInfo|TLLM_GEN_VERSION|CudaArch' "$P/fmhaKernels.h" "$P/fmhaKernels.cpp" | head -180
echo TARGET_ARCHIVE_SAMPLE
git ls-tree -r --name-only "$T" "$P/cubin" | grep 'Sm107a.*Context.cubin.tar.zst' | head -1 | while IFS= read -r f; do echo "$f"; git show "$T:$f" | head -c 180; echo; done
echo METADATA_REGISTRATION
git show "$T:$P/cubin/kernelMetaInfo.h" | grep -n -m3 -C3 'FmhaSm107aKernel_.*_cubin,' | head -35
echo TARGET_REFERENCE_LINES
git show "$T:cpp/cmake/modules/tllm_cubin_archive.cmake" | sed -n '100,119p'
git show "$T:$P/CMakeLists.txt" | nl -ba | sed -n '48,78p'Length of output: 31513
SEMANTIC_REVIEW
Verdict: INCONCLUSIVE. Git confirms all three supplied commits and the stated merge-base. The branches diverge. The head changes only the FMHA CMake file; the target leaves that file unchanged. The three-way preview cleanly combines the head’s SM107 retention change with the target’s SM107 metadata rows. I inspected architecture filtering, metadata registration, archive filenames, and runtime kernel lookup. At the target, all 4,006 SM107 metadata symbol names have matching archive filenames. At the merge-base and head, all 3,920 do. I found no source-supported metadata-to-archive incompatibility. Material limit: The target’s cubin archives are Git LFS pointers in the available checkout, not readable cubin payloads. The head’s change makes those SM107 payloads available in builds without architecture 107 enabled. Archive names and metadata cannot establish whether the target’s changed payloads load and execute correctly on that path. The payloads, or a build and runtime verification using these fixed revisions, are needed for that judgment. This is not a claim that a payload is defective. SEMANTIC_RESULT request_id=3a6f174e-90cf-464d-bc4a-fbc730b39456 head=a709ddfe0c30e7f1f0625d1d7f2cfb05982e0b25 target=af131fc340c62e77f5e5238f31c9bb7cb4364307 merge_base=b91158667191d2226006e6b20557bcba2d02d2a1 verdict=INCONCLUSIVE You are interacting with an AI system. |
…HA kernels On Rubin (SM107), context attention fell back to the unfused path and tried to allocate a 97.6 GiB workspace, OOMing and then faulting. Release builds do not enable 107, which dropped the Sm107a trtllm-gen cubins and their kernelMetaInfo.h rows. Sm100f has no context kernel, so Rubin was left with none at all. Omit 107 from the prune list and clear EXCLUDE_SM_107 for this directory only. The cubins are prebuilt and loaded via cuModuleLoadData, so no SM107 device codegen is added. Signed-off-by: Chulian Zhang <851104+zhangcl@users.noreply.github.com>
a709ddf to
ffa65a9
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #75681 [ run ] triggered by Bot. Commit: |
|
Evaluate behavioral incompatibilities involving the supplied PR changes and the Verify the three full commit IDs and their merge-base. For divergent branches, When head contains target ( First identify changed or removed interfaces and symbols, then search repository Classify findings in prose as cross-branch interactions or PR-local compatibility Use read-only source and Git inspection. Do not modify checked-out files, execute Assess independent findings separately. Uncertainty about one path does not
Write exactly one standalone heading
Then give the verdict, findings or inspected paths, and limitations. For PASS and Copy the supplied identity and commit IDs verbatim. Always include the result Analyze only these fixed revisions in NVIDIA/TensorRT-LLM, not the hosting PR: |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -u
H=ffa65a928521def383d8754377cea3daf3add2e3
T=ae3531adf3edcd833844559ecb791ab29b398d5d
B=ae3531adf3edcd833844559ecb791ab29b398d5d
for x in "$H" "$T" "$B"; do printf 'commit %s: ' "$x"; git rev-parse --verify "$x^{commit}" || :; done
printf 'actual merge-base: '; git merge-base "$H" "$T" || :
printf 'repository: '; git remote -v | head -2
echo 'diff stats'; git diff --stat "$T" "$H" | tail -70
echo 'changed file names'; git diff --name-status "$T" "$H" | head -160
echo 'head CMake diff'; git diff "$T" "$H" -- cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/CMakeLists.txtLength of output: 2894 🏁 Script executed: #!/bin/bash
set -u
H=ffa65a928521def383d8754377cea3daf3add2e3
T=ae3531adf3edcd833844559ecb791ab29b398d5d
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo 'HEAD_CMAKE'; git show "$H:$P/CMakeLists.txt" | nl -ba | head -110
echo 'TARGET_CMAKE'; git show "$T:$P/CMakeLists.txt" | nl -ba | head -105
echo 'CMAKE_HELPER'; git show "$T:cpp/cmake/modules/tllm_cubin_archive.cmake" | nl -ba | head -230
echo 'FILTER_FUNCTION'; git grep -n -A100 -B8 'function(filter_source_cuda_architectures' "$T" -- cpp/cmake
echo 'EXCLUDE_REFERENCES'; git grep -n -E 'EXCLUDE_SM_107|EXCLUDE_SM_100|EXCLUDE_SM_103|EXCLUDE_SM_100F' "$T" -- 'cpp/*.txt' 'cpp/*.cmake' 'cpp/tensorrt_llm/kernels/*' 'cpp/cmake/*' | head -100
echo 'FMHA_TREE'; git ls-tree --name-only "$T:$P" | head -90
echo 'META_USAGE'; git grep -n -E 'kernelMetaInfo.h|Sm107a|SM_107|sm_107' "$T" -- "$P" ':!*.cubin.tar.zst' | head -100Length of output: 70834 🏁 Script executed: #!/bin/bash
set -u
T=ae3531adf3edcd833844559ecb791ab29b398d5d
H=ffa65a928521def383d8754377cea3daf3add2e3
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo 'FILTER_TAIL'; git show "$T:cpp/cmake/modules/cuda_configuration.cmake" | nl -ba | sed -n '687,711p'
echo 'ALL_MACRO_DEFINITION_REFERENCES'; git grep -n -E 'EXCLUDE_SM_107|filter_source_cuda_architectures\(' "$T" -- '*.cmake' '*CMakeLists.txt' | head -115
echo 'FMHA_CPP_CU_LIST'; git ls-tree -r --name-only "$T" "$P" | grep -E '\.(cpp|cu)$' | head -95
echo 'HEAD_SM107_FILES'; git ls-tree -r --name-only "$H" "$P" | grep -Ei 'sm_?107' | grep -v '\.cubin\.tar\.zst$' | head -45
echo 'HEADER_INCLUDES'; git grep -n -F 'kernelMetaInfo.h' "$T" -- 'cpp/*' | head -65
echo 'RUNTIME_REFERENCES'; git grep -n -E 'FmhaSm107|Sm107a|mKernelMetaInfo|kernelMetaInfo|TLLM_RUBIN_FEATURES' "$T" -- "$P/fmhaRunner.cpp" "$P/fmhaKernels.h" "$P/trtllmGen_fmha_export" | head -90
echo 'AST_OUTLINE_RUNNER'; ast-grep outline "$P/fmhaRunner.cpp" --type class,function,method --match 'Fmha|fmha|Kernel|kernel|run' | head -70
echo 'ARCHIVE_META_COUNTS'; python - <<'PY'
import subprocess,re,collections
T='ae3531adf3edcd833844559ecb791ab29b398d5d'
P='cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/'
text=subprocess.check_output(['git','show',f'{T}:{P}kernelMetaInfo.h'],text=True)
paths=subprocess.check_output(['git','ls-tree','-r','--name-only',T,P],text=True).splitlines()
archives={p.rsplit('/',1)[-1].removesuffix('.cubin.tar.zst') for p in paths if p.endswith('.cubin.tar.zst')}
syms=set(re.findall(r'\b(FmhaSm107aKernel_\w+)_cubin\b',text))
a={s for s in archives if s.startswith('FmhaSm107aKernel_')}
print('archives',len(a),'referenced SM107',len(syms),'missing archives',len(syms-a),'unreferenced archives',len(a-syms))
print('examples missing',sorted(syms-a)[:3])
print('examples unreferenced',sorted(a-syms)[:3])
PYLength of output: 15710 🏁 Script executed: #!/bin/bash
set -u
T=ae3531adf3edcd833844559ecb791ab29b398d5d
H=ffa65a928521def383d8754377cea3daf3add2e3
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo 'EXCLUDE_ALL_TEXT'; git grep -n -F 'EXCLUDE_SM_107' "$T" -- '*.txt' '*.cmake' '*.h' '*.cpp' '*.cu' | grep -v 'kernelMetaInfo.h' | head -75
echo 'ARCH_CONFIG_FIRST'; git show "$T:cpp/cmake/modules/cuda_configuration.cmake" | nl -ba | sed -n '72,215p'
echo 'PARENT_CMAKE_REFS'; git grep -n -E 'add_definitions|EXCLUDE_SM_|add_subdirectory\(.*(trtllmGen|fmha)' "$T" -- cpp/CMakeLists.txt cpp/tensorrt_llm/CMakeLists.txt cpp/tensorrt_llm/kernels/CMakeLists.txt cpp/tensorrt_llm/kernels/trtllmGenKernels/CMakeLists.txt | head -120
echo 'META_HEADER_TOP'; git show "$T:$P/cubin/kernelMetaInfo.h" | nl -ba | sed -n '17,42p;7078,7091p;15087,15122p;18668,18690p;22678,22706p'
echo 'KERNEL_SELECTION'; git show "$T:$P/fmhaKernels.h" | nl -ba | sed -n '30,70p;1391,1445p;1550,1610p'
echo 'RUNNER_START'; git show "$T:$P/fmhaRunner.cpp" | nl -ba | sed -n '30,111p'
echo 'BLOB_EXAMPLE'; f=$(git ls-tree -r --name-only "$T" "$P/cubin" | grep 'FmhaSm107aKernel_.*Context.cubin.tar.zst$' | head -1); echo "$f"; git cat-file -s "$T:$f"; git show "$T:$f" | head -c 155; echo
echo 'HEAD_TARGET_HEADER_IDENTICAL'; for c in "$T" "$H"; do git rev-parse "$c:$P/cubin/kernelMetaInfo.h"; doneLength of output: 47031 🏁 Script executed: #!/bin/bash
set -u
T=ae3531adf3edcd833844559ecb791ab29b398d5d
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo 'CM_EXCLUSION_ALL'; git grep -n -E 'EXCLUDE_SM|PLACEHOLDER_KERNELS|CMAKE_CUDA_ARCHITECTURES_ORIG|CMAKE_CUDA_ARCHITECTURES_FAMILIES' "$T" -- 'cpp/*CMakeLists.txt' 'cpp/**/*.cmake' 'cpp/**/CMakeLists.txt' | head -135
echo 'CMAKE_FILE_CPP'; git show "$T:cpp/CMakeLists.txt" | nl -ba | sed -n '245,320p;370,445p'
echo 'KERNELS_PARENT'; git show "$T:cpp/tensorrt_llm/kernels/trtllmGenKernels/CMakeLists.txt" | nl -ba | head -100
echo 'FMHA_LOOKUP_SYMBOLS'; git show "$T:$P/fmhaKernels.h" | nl -ba | sed -n '1100,1225p;1490,1545p'
echo 'GATE_REFERENCE_ALL'; git grep -n 'EXCLUDE_SM_100F' "$T" -- "$P" | grep -v kernelMetaInfo.h | head -35
echo 'SM107_ARCH_SUPPORTED'; git grep -n -F 'isSM100Family' "$T" -- cpp/tensorrt_llm/common | head -30
echo 'TRTLLM_GEN_FMHA_CALLERS'; git grep -n -E '\bTllmGenFmhaRunner\b' "$T" -- cpp/tensorrt_llm | head -45Length of output: 31511 🏁 Script executed: #!/bin/bash
set -u
T=ae3531adf3edcd833844559ecb791ab29b398d5d
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo 'ARCH_SETUP'; git show "$T:cpp/cmake/modules/cuda_configuration.cmake" | nl -ba | sed -n '324,484p'
echo 'CPP_ARCH_SETUP'; git grep -n -E 'setup_cuda_architectures|CMAKE_CUDA_ARCHITECTURES_RAW|CMAKE_CUDA_ARCHITECTURES[ =]' "$T" -- cpp/CMakeLists.txt | head -65
echo 'FMHA_FACTORY_SETUP'; git show "$T:$P/fmhaKernels.h" | nl -ba | sed -n '116,215p;1450,1490p'
echo 'FMHA_DISPATCH'; git show "$T:cpp/tensorrt_llm/kernels/fmhaDispatcher.cpp" | nl -ba | sed -n '32,128p'
echo 'ATTENTION_CALLER'; git show "$T:cpp/tensorrt_llm/common/attentionOp.cpp" | nl -ba | sed -n '3160,3220p'
echo 'FMHA_HELPER_AGG_TAIL'; git show "$T:cpp/cmake/modules/tllm_cubin_archive.cmake" | nl -ba | sed -n '225,270p'
echo 'LFS_POINTER_TOTAL'; git ls-tree -r --name-only "$T" "$P/cubin" | grep 'FmhaSm107aKernel_.*cubin.tar.zst$' | head -1 | while read -r f; do git cat-file -s "$T:$f"; doneLength of output: 24853
SEMANTIC_REVIEW
Verdict: INCONCLUSIVE. Git verifies all three supplied commit IDs. The merge-base is the supplied target, I found no source-proven incompatibility in the inspected build and lookup paths. The head retains SM 107 archives and clears Material limit: The inspected SM 107 archive in the fixed revision is a Git LFS pointer, not a readable cubin payload in the available checkout. Filename and metadata checks cannot verify that newly retained cubins load and execute on SM 107. That requires the payloads or build and runtime verification at these fixed revisions. No payload defect is established. SEMANTIC_RESULT request_id=3b7c10c3-4a03-4fcd-9c4a-0596e0f7de19 head=ffa65a928521def383d8754377cea3daf3add2e3 target=ae3531adf3edcd833844559ecb791ab29b398d5d merge_base=ae3531adf3edcd833844559ecb791ab29b398d5d verdict=INCONCLUSIVE You are interacting with an AI system. |
|
PR_Github #75681 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #75719 [ run ] triggered by Bot. Commit: |
|
PR_Github #75719 [ run ] completed with state
|
|
Evaluate behavioral incompatibilities involving the supplied PR changes and the Verify the three full commit IDs and their merge-base. For divergent branches, When head contains target ( First identify changed or removed interfaces and symbols, then search repository Classify findings in prose as cross-branch interactions or PR-local compatibility Use read-only source and Git inspection. Do not modify checked-out files, execute Assess independent findings separately. Uncertainty about one path does not
Write exactly one standalone heading
Then give the verdict, findings or inspected paths, and limitations. For PASS and Copy the supplied identity and commit IDs verbatim. Always include the result Analyze only these fixed revisions in NVIDIA/TensorRT-LLM, not the hosting PR: |
|
Tip For best results, initiate chat on the files or code changes. 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -u
H=ffa65a928521def383d8754377cea3daf3add2e3
T=bdd012579f8941ec188ea1db46ca3f77da0b9de0
B=ae3531adf3edcd833844559ecb791ab29b398d5d
for c in "$H" "$T" "$B"; do git rev-parse --verify "$c^{commit}"; git cat-file -t "$c"; done
echo ACTUAL_MERGE_BASE
git merge-base "$H" "$T"
echo HEAD_DIFF
git diff --stat "$B" "$H"
git diff "$B" "$H"
echo TARGET_DIFF_NAMES
git diff --name-status "$B" "$T"
echo TARGET_RELEVANT_DIFF
git diff "$B" "$T" -- cpp/cmake cpp/CMakeLists.txt cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/CMakeLists.txt cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/fmhaKernels.h cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/fmhaRunner.cpp
echo MERGE_PREVIEW
git merge-tree "$B" "$H" "$T" | head -180Length of output: 12294 🏁 Script executed: #!/bin/bash
set -u
H=ffa65a928521def383d8754377cea3daf3add2e3
T=bdd012579f8941ec188ea1db46ca3f77da0b9de0
B=ae3531adf3edcd833844559ecb791ab29b398d5d
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo OUTLINES
ast-grep outline cpp/cmake/modules/tllm_cubin_archive.cmake --items all | head -60
ast-grep outline "$P/fmhaKernels.h" --match 'load|Kernel|Factory' | head -100
ast-grep outline "$P/fmhaRunner.cpp" --items all | head -80
echo CMAKE
git show "$H:$P/CMakeLists.txt" | nl -ba | sed -n '18,110p'
echo ARCHIVE_HELPER
git show "$T:cpp/cmake/modules/tllm_cubin_archive.cmake" | nl -ba | sed -n '20,270p'
echo FILTER
git show "$T:cpp/cmake/modules/cuda_configuration.cmake" | nl -ba | sed -n '590,710p'
echo FIXED_REFERENCES
for c in "$H" "$T"; do
echo "$c"
git grep -n -E 'EXCLUDE_SM_107|filter_source_cuda_architectures\(|tllm_add_cubin_archive_sources\(' "$c" -- '*.cmake' '*CMakeLists.txt' | head -65
git grep -n -E 'TllmGenFmhaRunner|kernelMetaInfo.h' "$c" -- cpp/tensorrt_llm ':!**/cubin/kernelMetaInfo.h' | head -55
done
echo RELEVANT_BLOB_IDENTITIES
for c in "$B" "$H" "$T"; do
echo "$c"
git rev-parse "$c:$P/fmhaKernels.h" "$c:$P/cubin/kernelMetaInfo.h" "$c:cpp/cmake/modules/cuda_configuration.cmake" "$c:cpp/cmake/modules/tllm_cubin_archive.cmake"
done
echo TARGET_LINK_DIFF
git diff "$B" "$T" -- cpp/tensorrt_llm/CMakeLists.txt cpp/tensorrt_llm/common/CMakeLists.txt cpp/tensorrt_llm/runtime/utils/CMakeLists.txtLength of output: 48239 🏁 Script executed: #!/bin/bash
set -u
H=ffa65a928521def383d8754377cea3daf3add2e3
T=bdd012579f8941ec188ea1db46ca3f77da0b9de0
B=ae3531adf3edcd833844559ecb791ab29b398d5d
P=cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha
echo ARCH_DEFINITIONS
git show "$T:cpp/cmake/modules/cuda_configuration.cmake" | nl -ba | sed -n '324,485p'
echo KERNEL_FACTORY_AND_LOAD
git show "$T:$P/fmhaKernels.h" | nl -ba | sed -n '1390,1438p;1450,1610p'
echo RUNNER
git show "$T:$P/fmhaRunner.cpp" | nl -ba | sed -n '32,106p'
echo DISPATCH_CALLER
git show "$T:cpp/tensorrt_llm/kernels/fmhaDispatcher.cpp" | nl -ba | sed -n '32,126p'
echo ATTENTION_CALLER
git show "$T:cpp/tensorrt_llm/common/attentionOp.cpp" | nl -ba | sed -n '3160,3210p'
echo SOURCE_AND_TEST_REFERENCES
git ls-tree -r --name-only "$H" "$P" | grep -E '\.(cpp|cu)$' | head -30
git grep -n -E 'EXCLUDE_SM_|add_subdirectory\(fmha|trtllm_gen_fmha' "$T" -- cpp/CMakeLists.txt cpp/tensorrt_llm/kernels/CMakeLists.txt cpp/tensorrt_llm/kernels/trtllmGenKernels/CMakeLists.txt
git grep -n -E 'TllmGenFmhaRunner|Sm107a|SM107' "$T" -- cpp/tests tests -g 2>/dev/null | head -35
echo ARTIFACT_ANALYSIS
python - <<'PY'
import subprocess,re
H='ffa65a928521def383d8754377cea3daf3add2e3'
T='bdd012579f8941ec188ea1db46ca3f77da0b9de0'
B='ae3531adf3edcd833844559ecb791ab29b398d5d'
p='cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/'
for c in (B,H,T):
text=subprocess.check_output(['git','show',f'{c}:{p}kernelMetaInfo.h'],text=True)
paths=subprocess.check_output(['git','ls-tree','-r','--name-only',c,p],text=True).splitlines()
archives={x.rsplit('/',1)[-1].removesuffix('.cubin.tar.zst'):x for x in paths if x.endswith('.cubin.tar.zst')}
syms=set(re.findall(r'\b(FmhaSm107aKernel_\w+)_cubin\b',text))
print(c,'SM107 metadata symbols',len(syms),'matching archives',len(syms&archives.keys()),'missing',sorted(syms-archives.keys())[:5])
if c==T:
f=next(archives[s] for s in sorted(syms) if s.endswith('Context'))
blob=subprocess.check_output(['git','show',f'{c}:{f}'])
print('SAMPLE',f,'bytes',len(blob),'content',blob[:160])
lines=text.splitlines()
for i,line in enumerate(lines,1):
if 'EXCLUDE_SM_107' in line:
print('GATE',i,line)
print('\n'.join(f'{j+1}: {lines[j]}' for j in range(i,min(i+3,len(lines)))))
print('ARCHIVE_DIFF_BASE_TARGET')
print(subprocess.check_output(['git','diff','--stat',B,T,'--',p],text=True))
PY
echo PREVIEW_CONFLICTS
git merge-tree "$B" "$H" "$T" | grep -n -E '^changed in both|^<<<<<<<|^=======|^>>>>>>>' | head -25
echo TARGET_CMAKE_CITATION
git show "$T:$P/CMakeLists.txt" | nl -ba | sed -n '49,70p'Length of output: 43888
SEMANTIC_REVIEW
Verdict: INCONCLUSIVE. No concrete, in-scope behavioral incompatibility was established. Required cubin payloads were unavailable for inspection. Revision and integration checks Git verified all three supplied commit IDs. The actual merge-base is
Inspected contracts I checked architecture definitions, archive selection, generated symbol names, metadata gates, runner arguments, and runtime dispatch. References included unchanged callers in At all three revisions, all 4,006 SM107 metadata symbol names match archive filenames. The metadata, archive helper, architecture-filter helper, and The existing Material limitation The inspected SM107 context archive contains a Git LFS pointer, not the binary payload. Source and filename checks cannot establish whether newly retained payloads satisfy the metadata contract and load correctly on SM107. No payload defect is established. Please retry with the fixed-revision payloads available for read-only inspection. I did not execute project code or tests, and I did not use the hosting PR’s current revisions or CI discussion as evidence. SEMANTIC_RESULT request_id=ce01bed9-2f28-4c6d-b9f9-03d0ba24cfe2 head=ffa65a928521def383d8754377cea3daf3add2e3 target=bdd012579f8941ec188ea1db46ca3f77da0b9de0 merge_base=ae3531adf3edcd833844559ecb791ab29b398d5d verdict=INCONCLUSIVE You are interacting with an AI system. |
Description
On Rubin (SM107), context attention fell back to the unfused path and tried to allocate a 97.6 GiB workspace, OOMing and then faulting.
Release builds do not enable 107, which dropped the Sm107a trtllm-gen cubins and their
kernelMetaInfo.hrows.Sm100fhas no context kernel, so Rubin was left with none at all.Fix: omit 107 from the prune list and clear
EXCLUDE_SM_107for this directory only. The cubins are prebuilt and loaded viacuModuleLoadData, so no SM107 device codegen is added.Test Coverage
accuracy/test_llm_api_pytorch.py::TestNemotronV3Ultra::test_nvfp4_8gpus[attention_dp_on-cutedsl]Passes on 8x VR200 with a stock build; workspace 97.6 GiB -> ~198 MB.