[None][fix] Pass max_seq_len to the Qwen-VL attention helper in the MiniCPM-V 4.6 vision encoder - #19641
[None][fix] Pass max_seq_len to the Qwen-VL attention helper in the MiniCPM-V 4.6 vision encoder#19641David-Wu1119 wants to merge 1 commit into
Conversation
…iniCPM-V 4.6 vision encoder MiniCPMV4_6VisionModel._make_attn_metadata calls the shared _prepare_qwen_vl_vision_attn_metadata helper. NVIDIA#16568 made max_seq_len a required keyword-only argument of that helper and updated the Qwen2-VL and Qwen3-VL callers, but not this one, so every MiniCPM-V 4.6 image or video request raised TypeError in the vision encoder. Pass max(seq_lens), which is the value the helper used before NVIDIA#16568. The metadata here is built per call, so this keeps the vision encoder's behavior as it was when MiniCPM-V 4.6 support landed. The new unit test replaces the helper with one that has the same signature and checks the call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David-Wu1119 <133224895+David-Wu1119@users.noreply.github.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA/TensorRT-LLM/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. Walkthrough
ChangesVision attention metadata
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~5 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to This fixes a crash in MiniCPM-V 4.6 image/video requests by passing a required argument to a shared helper; the fix is small, targeted, and tested. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Description
MiniCPMV4_6VisionModel._make_attn_metadatabuilds the vision encoder's attention metadata with the shared Qwen-VL helper:MiniCPM-V 4.6 support (#15976, 2026-07-28) was written against the helper's original
(seq_lens, attn_metadata)signature. A week later #16568 mademax_seq_lena required keyword-only argument and updated the Qwen2-VL and Qwen3-VL callers, but not this one. Since then, every MiniCPM-V 4.6 image or video request fails in the vision encoder, both before and after the ViT merger layer:The fix passes
max_seq_len=max(seq_lens). Before #16568, the helper set exactly this value (attn_metadata.max_seq_len = max(seq_lens)), so the vision encoder behaves as it did when MiniCPM-V 4.6 support landed. I kept the pre-#16568 value instead of picking a fixed capacity. This metadata is sized per call (max_num_tokens=sum(seq_lens) + 1), and I didn't want to guess a capacity for MiniCPM. If you want MiniCPM to get the same stable attention-op cache key that #16568 gave Qwen, a fixed value can be a follow-up.Test Coverage
test_vision_attn_metadata_passes_max_seq_lenintests/unittest/_torch/modeling/test_modeling_minicpmv4_6.py, which is already inl0_a10. It follows the existingtest_qwen3_vision_prepare_metadata_passes_fixed_max_seq_len: it replaces the helper with one that has the same signature, calls_make_attn_metadata, and checks the arguments. Onmainthe call raises the TypeError above; with this change it passes.tensorrt_llmlocally. Instead I pulled_make_attn_metadataand the real_prepare_qwen_vl_vision_attn_metadataout of the source and ran them with a stub metadata class:pre-commit runpasses on both files (ruff, ruff-format, codespell, test-list checks).PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
I found this with pylint's
missing-kwoa(E1125) check. The change was written with AI assistance and I reviewed every line.🤖 Generated with Claude Code