Skip to content

[None][fix] Pass max_seq_len to the Qwen-VL attention helper in the MiniCPM-V 4.6 vision encoder - #19641

Open
David-Wu1119 wants to merge 1 commit into
NVIDIA:mainfrom
David-Wu1119:fix/minicpmv4-6-vision-max-seq-len
Open

David-Wu1119 wants to merge 1 commit into
NVIDIA:mainfrom
David-Wu1119:fix/minicpmv4-6-vision-max-seq-len

Conversation

@David-Wu1119

@David-Wu1119 David-Wu1119 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Description

MiniCPMV4_6VisionModel._make_attn_metadata builds the vision encoder's attention metadata with the shared Qwen-VL helper:

return _prepare_qwen_vl_vision_attn_metadata(seq_lens, attn_metadata)

MiniCPM-V 4.6 support (#15976, 2026-07-28) was written against the helper's original (seq_lens, attn_metadata) signature. A week later #16568 made max_seq_len a required keyword-only argument and updated the Qwen2-VL and Qwen3-VL callers, but not this one. Since then, every MiniCPM-V 4.6 image or video request fails in the vision encoder, both before and after the ViT merger layer:

TypeError: _prepare_qwen_vl_vision_attn_metadata() missing 1 required keyword-only argument: 'max_seq_len'

The fix passes max_seq_len=max(seq_lens). Before #16568, the helper set exactly this value (attn_metadata.max_seq_len = max(seq_lens)), so the vision encoder behaves as it did when MiniCPM-V 4.6 support landed. I kept the pre-#16568 value instead of picking a fixed capacity. This metadata is sized per call (max_num_tokens=sum(seq_lens) + 1), and I didn't want to guess a capacity for MiniCPM. If you want MiniCPM to get the same stable attention-op cache key that #16568 gave Qwen, a fixed value can be a follow-up.

Test Coverage

  • New test_vision_attn_metadata_passes_max_seq_len in tests/unittest/_torch/modeling/test_modeling_minicpmv4_6.py, which is already in l0_a10. It follows the existing test_qwen3_vision_prepare_metadata_passes_fixed_max_seq_len: it replaces the helper with one that has the same signature, calls _make_attn_metadata, and checks the arguments. On main the call raises the TypeError above; with this change it passes.
  • I don't have a CUDA machine, so I couldn't import tensorrt_llm locally. Instead I pulled _make_attn_metadata and the real _prepare_qwen_vl_vision_attn_metadata out of the source and ran them with a stub metadata class:
main  -> TypeError: _prepare_qwen_vl_vision_attn_metadata() missing 1 required keyword-only argument: 'max_seq_len'
fix   -> max_seq_len = 64, cu_q_seqlens = [0, 64, 80], prepare_encoder_only() called
  • pre-commit run passes on both files (ruff, ruff-format, codespell, test-list checks).

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

I found this with pylint's missing-kwoa (E1125) check. The change was written with AI assistance and I reviewed every line.

🤖 Generated with Claude Code

…iniCPM-V 4.6 vision encoder

MiniCPMV4_6VisionModel._make_attn_metadata calls the shared
_prepare_qwen_vl_vision_attn_metadata helper. NVIDIA#16568 made max_seq_len a
required keyword-only argument of that helper and updated the Qwen2-VL and
Qwen3-VL callers, but not this one, so every MiniCPM-V 4.6 image or video
request raised TypeError in the vision encoder.

Pass max(seq_lens), which is the value the helper used before NVIDIA#16568. The
metadata here is built per call, so this keeps the vision encoder's behavior
as it was when MiniCPM-V 4.6 support landed. The new unit test replaces the
helper with one that has the same signature and checks the call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: David-Wu1119 <133224895+David-Wu1119@users.noreply.github.com>
@David-Wu1119
David-Wu1119 requested a review from a team as a code owner September 26, 2026 20:50
@David-Wu1119
David-Wu1119 requested review from 2ez4bz and Wanli-Jiang and a lite review from Copilot September 26, 2026 20:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/TensorRT-LLM/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: d62d80f5-c2f6-4168-9339-aadc7ccd8761

📥 Commits

Reviewing files that changed from the base of the PR and between b88149e and 2d49523.

📒 Files selected for processing (2)
  • tensorrt_llm/_torch/models/modeling_minicpmv4_6.py
  • tests/unittest/_torch/modeling/test_modeling_minicpmv4_6.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

MiniCPMV4_6VisionModel._make_attn_metadata now passes the maximum sequence length to the vision attention metadata helper. A unit test checks that the helper receives the sequence lengths and their maximum.

Changes

Vision attention metadata

Layer / File(s) Summary
Pass and verify maximum sequence length
tensorrt_llm/_torch/models/modeling_minicpmv4_6.py, tests/unittest/_torch/modeling/test_modeling_minicpmv4_6.py
_make_attn_metadata passes max(seq_lens) as max_seq_len. The test checks the helper arguments for sequence lengths [64, 16].

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~5 minutes

Change: Bug fix

Suggested reviewers: bowenfu

Merge Risk: ⚪ Minimal · up to 2d495

This fixes a crash in MiniCPM-V 4.6 image/video requests by passing a required argument to a shared helper; the fix is small, targeted, and tested.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the fix, the affected MiniCPM-V 4.6 vision encoder, and the missing max_seq_len argument. It follows the required [None][fix] format.
Description check ✅ Passed The description explains the issue, the fix, test coverage, validation results, and checklist status. It provides sufficient context and is largely complete for this change.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@svc-trtllm-gh-bot svc-trtllm-gh-bot added the Community want to contribute PRs initiated from Community label Sep 26, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Community want to contribute PRs initiated from Community

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants