Skip to content

optimize qwen3.8-flash-next - #185

Draft
hjh0119 wants to merge 12 commits into
modelscope:mainfrom
hjh0119:qwen4-0903
Draft

optimize qwen3.8-flash-next#185
hjh0119 wants to merge 12 commits into
modelscope:mainfrom
hjh0119:qwen4-0903

Conversation

@hjh0119

@hjh0119 hjh0119 commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

hjh0119 added 12 commits August 26, 2026 17:23
main landed the same feature independently in modelscope#174, so qwen4_exp.py /
ple.py / qsa_indexer.py / hyper_connection_gated.py came out as add/add
conflicts. Resolved to this branch's versions -- they are supersets:
  select_mask              -> selection_as_mask (+ selection_as_token_indices,
                              select_token_indices_thd for the sparse kernel)
  _set_ple_ngram_embedding -> fill_table_from_hf / export_table_to_hf
                              (offload-aware, and reduces the offload flag
                              across pp before gating pp collectives)
  _warn_qsa_fallback_once  -> dropped with the fallback path itself

Kept main's ple_seed instead of this branch's hardcoded _PLE_SEED: the
parser derives it from text_config.seed (defaulting to 1234), which stays
configurable and still avoids the vLLM config-pollution issue. Wired it
through Qwen4ExpTextNGramEmbedding with the same fail-loud treatment as
eos_token_id / split_ngram_parts.

Also picks up modelscope#175 (packed sequence length handling for mcore 0.16.0/0.16.1).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant