[feat] Add fastvideo serve configs for Wan CUDA models - #1801
Conversation
Merge Protections🟠 1 of 1 protections blocking · waiting on 🤖 CI
🟠 PR merge requirementsWaiting for
Waiting checks:
|
There was a problem hiding this comment.
🟡 Changes recommended
Two of the new serving YAMLs have header comments that instruct users to send inputs.image_path, which does not match the OpenAI-compatible request schema supported by the server.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR adds fastvideo serve (OpenAI-compatible) example configs for four Wan CUDA Diffusers checkpoints so they can be launched as a persistent server similarly to the existing MiniMax H3 serving examples.
Changes:
- Added serving config for FastWan2.1 T2V 1.3B (1 GPU).
- Added serving configs for Wan2.2 T2V A14B (2 GPUs) and Wan2.2 TI2V 5B (1 GPU).
- Added serving config for Wan2.1 I2V 14B 480P (2 GPUs) including I2V-specific pipeline selection.
File summaries
| File | Description |
|---|---|
| examples/serving/openai_fastwan21_1_3b.yaml | Adds an OpenAI-compatible serving config for FastWan2.1 T2V 1.3B (single GPU, compiled). |
| examples/serving/openai_wan21_i2v_14b.yaml | Adds an OpenAI-compatible serving config for Wan2.1 I2V 14B 480P (2 GPUs, I2V workload). |
| examples/serving/openai_wan22_t2v_a14b.yaml | Adds an OpenAI-compatible serving config for Wan2.2 T2V A14B (2 GPUs). |
| examples/serving/openai_wan22_ti2v_5b.yaml | Adds an OpenAI-compatible serving config for Wan2.2 TI2V 5B (single GPU, supports image-conditional requests). |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| # OpenAI-compatible server for Wan2.2 TI2V 5B. Supports both text-to-video | ||
| # (omit inputs.image_path) and image-to-video (set it) from the same model. |
| # OpenAI-compatible server for Wan2.1 I2V 14B 480P. Each request must supply | ||
| # its own source image (POST inputs.image_path); there is no default image. |
Review:
|
| config | gpus | tp/sp | sampling | steps | guidance |
|---|---|---|---|---|---|
| fastwan21_1_3b | 1 | 1/1 | 480x832, 81f, fps16 | 3 | 3.0 |
| wan21_i2v_14b | 2 | 2/2 | 480x832, 77f, fps16 | 40 | 5.0 |
| wan22_t2v_a14b | 2 | 1/2 | 720x1280, 81f, fps16 | 40 | 4.0/3.0 |
| wan22_ti2v_5b | 1 | 1/1 | 704x1280, 121f, fps24 | 50 | 5.0 |
Runtime verification is not possible in this sandbox (no GPU, no weights); the author already notes 3 of the 4 configs were not run. The 2-GPU configs and the i2v request path remain unverified end to end.
6540e9d to
dc6d734
Compare
Re-review at
|
SolitaryThinker
left a comment
There was a problem hiding this comment.
Reviewed and verified: the four Wan serve configs parse against current main and resolve to the same FastVideoArgs/SamplingParam values as their source generate configs. Rebased onto main and pushed fixes for the request-field docs, explicit parallelism, README coverage, and a parse test over examples/serving/openai_*.yaml. LGTM.
Wan could only be run as a one-off script — there was no way to run it as a server (
fastvideo serve), unlike MiniMax H3. This adds serving configs for all four Wan CUDA models, so they can be served the same way.openai_fastwan21_1_3b.yaml— FastWan2.1 1.3B, 1 GPUopenai_wan22_t2v_a14b.yaml— Wan2.2 A14B, 2 GPUsopenai_wan21_i2v_14b.yaml— Wan2.1 14B image-to-video, 2 GPUsopenai_wan22_ti2v_5b.yaml— Wan2.2 TI2V 5B, 1 GPUEvery setting in these configs is copied directly from the existing, working script/config for that model — nothing invented.
Test plan
openai_fastwan21_1_3b.yaml— ran for real: started the server, sent it a request, got back a real video (correct resolution and duration, matching the config).