Skip to content

Fix packed sequence length handling for Megatron-Core 0.16.0/0.16.1 - #175

Merged
tastelikefeet merged 4 commits into
modelscope:mainfrom
hazelduan:gdn_model_hybrid
Aug 29, 2026
Merged

Fix packed sequence length handling for Megatron-Core 0.16.0/0.16.1#175
tastelikefeet merged 4 commits into
modelscope:mainfrom
hazelduan:gdn_model_hybrid

Conversation

@hazelduan

Copy link
Copy Markdown
Contributor

Why

In Megatron-Core 0.16.0/0.16.1, PackedSeqParams.max_seqlen_q is defined as a Python int. The Qwen3.5 and Qwen3-Next THD packed-sequence paths called .item() on this value, causing:
AttributeError: 'int' object has no attribute 'item'

Changes

Convert max_seqlen_q with int(...) before allocating hidden states and attention masks. This supports both Python integers and scalar tensors without changing numerical computation.

Validation

  • Qwen3.5 training ran successfully for multiple steps on A3 with Megatron-Core 0.16.1.
  • Python compilation checks passed for both modified files.
  • Python int and scalar Tensor conversion cases passed.

GPU Compatibility

This change is device-agnostic and remains compatible with GPU training. max_seqlen_q is sequence-length metadata used only to construct tensor shapes; converting it with int(...) works with both Python integers and single-element CPU/CUDA/NPU tensors.

Experiments

截屏2026-08-29 下午4 10 28 Qwen3.5 training loss decreased and converged as expected. The change is also compatible with Megatron-Core 0.16.0–0.19.x, where `max_seqlen_q` remains integer metadata. Using `int(...)` supports both Python integers and scalar tensors without affecting model computation.

@tastelikefeet
tastelikefeet merged commit 4f2a95c into modelscope:main Aug 29, 2026
1 check passed
hjh0119 added a commit to hjh0119/mcore-bridge that referenced this pull request Sep 3, 2026
main landed the same feature independently in modelscope#174, so qwen4_exp.py /
ple.py / qsa_indexer.py / hyper_connection_gated.py came out as add/add
conflicts. Resolved to this branch's versions -- they are supersets:
  select_mask              -> selection_as_mask (+ selection_as_token_indices,
                              select_token_indices_thd for the sparse kernel)
  _set_ple_ngram_embedding -> fill_table_from_hf / export_table_to_hf
                              (offload-aware, and reduces the offload flag
                              across pp before gating pp collectives)
  _warn_qsa_fallback_once  -> dropped with the fallback path itself

Kept main's ple_seed instead of this branch's hardcoded _PLE_SEED: the
parser derives it from text_config.seed (defaulting to 1234), which stays
configurable and still avoids the vLLM config-pollution issue. Wired it
through Qwen4ExpTextNGramEmbedding with the same fail-loud treatment as
eos_token_id / split_ngram_parts.

Also picks up modelscope#175 (packed sequence length handling for mcore 0.16.0/0.16.1).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants