Skip to content

Add weight sharing and support for AMD (Vitis AI) - #1305

Open
Supreet Singh Palne (spalne) wants to merge 5 commits into
mainfrom
user/spalne/NPU
Open

Add weight sharing and support for AMD (Vitis AI)#1305
Supreet Singh Palne (spalne) wants to merge 5 commits into
mainfrom
user/spalne/NPU

Conversation

@spalne

Copy link
Copy Markdown
Contributor

Weight sharing for the decoder — Enables the Qwen3 decoder's context (prefill) and iterator (decode) stages to be compiled through a single shared EP context. Since both stages derive from the same weights, they now emit one shared weight .bin referenced by both graphs instead of duplicating the weights per stage reducing on disk bundle size and load-time memory. The grouping logic is execution-provider-agnostic, so it applies automatically to QNN, VitisAI, and OpenVINO.

Adds VitisAI as a supported execution provider for the Qwen3 transformer (context/iterator) stages, alongside the existing QNN. Building with now emits the correct per-stage session_options into genai_config.json, routing inference to the AMD NPU via the waic_target_vaiml_cpp_me VAIML C++ backend.

@spalne
Supreet Singh Palne (spalne) marked this pull request as ready for review August 13, 2026 15:54
@spalne
Supreet Singh Palne (spalne) requested a review from a team as a code owner August 13, 2026 15:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant