Skip to content

[feat] Add fastvideo serve configs for Wan CUDA models - #1801

Merged
SolitaryThinker merged 2 commits into
hao-ai-lab:mainfrom
Ishxn20:feat/wan-cuda-serving
Sep 15, 2026
Merged

SolitaryThinker merged 2 commits into
hao-ai-lab:mainfrom
Ishxn20:feat/wan-cuda-serving

Conversation

@Ishxn20

@Ishxn20 Ishxn20 commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Wan could only be run as a one-off script — there was no way to run it as a server (fastvideo serve), unlike MiniMax H3. This adds serving configs for all four Wan CUDA models, so they can be served the same way.

  • openai_fastwan21_1_3b.yaml — FastWan2.1 1.3B, 1 GPU
  • openai_wan22_t2v_a14b.yaml — Wan2.2 A14B, 2 GPUs
  • openai_wan21_i2v_14b.yaml — Wan2.1 14B image-to-video, 2 GPUs
  • openai_wan22_ti2v_5b.yaml — Wan2.2 TI2V 5B, 1 GPU

Every setting in these configs is copied directly from the existing, working script/config for that model — nothing invented.

Test plan

  • openai_fastwan21_1_3b.yaml — ran for real: started the server, sent it a request, got back a real video (correct resolution and duration, matching the config).
  • The other three configs are written the same careful way but not personally run — they need 2 GPUs, which wasn't available for testing.

Copilot AI lite review requested due to automatic review settings September 1, 2026 03:13
@mergify mergify Bot added the type: feat New feature or capability label Sep 1, 2026
@mergify

mergify Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Merge Protections

🟠 1 of 1 protections blocking · waiting on 🤖 CI

Protection Waiting on
🟠 PR merge requirements 🤖 CI

🟠 PR merge requirements

Waiting for

  • check-success=fastcheck-passed
  • check-success=full-suite-passed
Waiting checks: fastcheck-passed, full-suite-passed.
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • #approved-reviews-by>=1
  • check-success~=pre-commit
  • title~=(?i)^\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\]

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Two of the new serving YAMLs have header comments that instruct users to send inputs.image_path, which does not match the OpenAI-compatible request schema supported by the server.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR adds fastvideo serve (OpenAI-compatible) example configs for four Wan CUDA Diffusers checkpoints so they can be launched as a persistent server similarly to the existing MiniMax H3 serving examples.

Changes:

  • Added serving config for FastWan2.1 T2V 1.3B (1 GPU).
  • Added serving configs for Wan2.2 T2V A14B (2 GPUs) and Wan2.2 TI2V 5B (1 GPU).
  • Added serving config for Wan2.1 I2V 14B 480P (2 GPUs) including I2V-specific pipeline selection.
File summaries
File Description
examples/serving/openai_fastwan21_1_3b.yaml Adds an OpenAI-compatible serving config for FastWan2.1 T2V 1.3B (single GPU, compiled).
examples/serving/openai_wan21_i2v_14b.yaml Adds an OpenAI-compatible serving config for Wan2.1 I2V 14B 480P (2 GPUs, I2V workload).
examples/serving/openai_wan22_t2v_a14b.yaml Adds an OpenAI-compatible serving config for Wan2.2 T2V A14B (2 GPUs).
examples/serving/openai_wan22_ti2v_5b.yaml Adds an OpenAI-compatible serving config for Wan2.2 TI2V 5B (single GPU, supports image-conditional requests).
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1 to +2
# OpenAI-compatible server for Wan2.2 TI2V 5B. Supports both text-to-video
# (omit inputs.image_path) and image-to-video (set it) from the same model.
Comment on lines +1 to +2
# OpenAI-compatible server for Wan2.1 I2V 14B 480P. Each request must supply
# its own source image (POST inputs.image_path); there is no default image.
@SolitaryThinker

Copy link
Copy Markdown
Collaborator

Review: [feat] Add fastvideo serve configs for Wan CUDA models

Reviewed the four new YAMLs against the current main serve/request schema, the source generate configs/examples, and the Wan presets. All four parse cleanly and resolve to the same FastVideoArgs / SamplingParam values as their sources. Findings below, ordered by severity.

1. [P1] The two image-conditioned configs document a request field that does not exist

examples/serving/openai_wan21_i2v_14b.yaml:2, examples/serving/openai_wan22_ti2v_5b.yaml:2

inputs.image_path is FastVideo's internal typed-request path, not an OpenAI-compatible request field. VideoGenerationRequest is extra="forbid" (fastvideo/entrypoints/openai/protocol.py:120), so POST /v1/videos {"inputs": {"image_path": ...}} is rejected at admission with HTTP 400 Invalid request body: Extra inputs are not permitted. The server accepts input_reference / reference_url (path or URL string) or image_reference: [{"image_url": ...}]; request_adapter._image_sources maps those to image_path internally, and docs/design/server_contracts/openai.md documents the same. Copilot flagged this on the PR; it is still unaddressed.

2. [P2] examples/serving/README.md is now stale

examples/serving/README.md:5 says "The two FastH3 configs here are source-backed examples"; after this PR there are six configs. It also does not tell users the new served_model_name aliases (fastwan21-1.3b, wan21-i2v-14b, wan22-t2v-a14b, wan22-ti2v-5b) that must be passed as model / FASTVIDEO_MODEL, or how to send an image for the two image-conditioned configs. A short "Wan configs" section would close the loop.

3. [P2] Nothing parses the new configs in CI

No test under fastvideo/tests/ or tests/ loads examples/serving/*.yaml. A parametrized test over the directory using build_serve_config would catch schema drift cheaply (the recent lazy_module_load addition shows the schema does move). As written, these files get no automated validation.

4. [P3] CI is red on an unrelated lane

buildkite/pr-fastcheck/microscope-encoder-tests failed on the PR head. The diff only adds YAML under examples/, .github/scripts/plan_merge_ci.py treats examples/** as safe, and the base tree (a28f2bab4, PR #1794) passed the encoder lane. This looks like a flake/infra failure; re-run the build to unblock fastcheck-passed. I could not read the Buildkite log (the API requires auth).

5. [P3] Implicit parallelism on the A14B / TI2V configs

openai_wan22_t2v_a14b.yaml and openai_wan22_ti2v_5b.yaml omit engine.parallelism, relying on sp_size=-1 → num_gpus. It resolves correctly (A14B: tp=1/sp=2; TI2V: tp=1/sp=1), but the other two new configs and the FastH3 configs set it explicitly. Consider making it explicit for consistency.

6. [P3] Redundant settings (optional)

default_request.output.return_frames: false is a no-op — build_generation_request unconditionally sets return_frames: False (request_adapter.py:409). flow_shift: 3.0 (i2v) and dmd_denoising_steps: [1000, 757, 522] (1.3B) equal the pipeline-config defaults. Harmless and consistent with the source generate configs, so fine to keep.

Notes

  • offload.dit: true on the i2v/A14B/TI2V configs resolves to layerwise offload, not full CPU offload: dit_layerwise defaults to True (fastvideo/api/schema.py:28), and finalize_device_offload_policy disables dit_cpu_offload when layerwise is on. This matches the source generate configs/examples, but differs from the FastH3 serving configs, which explicitly set dit_layerwise: false. Worth knowing if full offload was the intent.
  • The 1.3B config sets offload.dit: false but leaves dit_layerwise at its default True, so the DiT still runs layerwise-offloaded (same as scripts/inference/inference_wan_VSA_DMD_1_3B.yaml). If "no DiT offload" was intended, set dit_layerwise: false.

Verification

Parsed all four configs with build_serve_config against current main (all OK) and resolved generator_config_to_fastvideo_args + SamplingParam.from_pretrained; effective values match the source configs/examples:

config gpus tp/sp sampling steps guidance
fastwan21_1_3b 1 1/1 480x832, 81f, fps16 3 3.0
wan21_i2v_14b 2 2/2 480x832, 77f, fps16 40 5.0
wan22_t2v_a14b 2 1/2 720x1280, 81f, fps16 40 4.0/3.0
wan22_ti2v_5b 1 1/1 704x1280, 121f, fps24 50 5.0

Runtime verification is not possible in this sandbox (no GPU, no weights); the author already notes 3 of the 4 configs were not run. The 2-GPU configs and the i2v request path remain unverified end to end.

@mergify mergify Bot added the scope: infra CI, tests, Docker, build label Sep 15, 2026
@SolitaryThinker

Copy link
Copy Markdown
Collaborator

Re-review at dc6d734f718d1206021a160e6fbeb21b408c9091

No new blocking findings from static review of all six changed files. This revision addresses the actionable points in the previous review:

  • Both image-conditioned YAML headers now use supported input_reference / image_reference fields; the request adapter maps them to internal image_path.
  • The README lists the Wan aliases, GPU counts, VSA launch requirement, and image-reference usage.
  • A14B and TI2V parallelism is explicit.
  • test_serve_examples.py discovers the OpenAI serving YAMLs and parses them through build_serve_config. .buildkite/scripts/unit_test.sh includes the entire fastvideo/tests/api/ directory, so this test is covered by the unit Fastcheck lane.

Compared the settings with the two source inference YAMLs and basic_wan2_2.py / basic_wan2_2_ti2v.py; no unintended recipe changes found. The earlier optional notes about inherited layerwise offload and redundant return_frames: false still apply, but are not blockers.

Verification: Python syntax compilation and whitespace checks passed. Could not run pytest or generation locally because this Python environment lacks PyTorch. The parse test does not establish GPU memory fit, generation quality, or repeated-request correctness; the three author-unverified configs still need runtime smoke evidence. Also, the PR test-plan explanation should distinguish the untested 1-GPU TI2V config from the two 2-GPU configs.

Current CI: pre-commit passes; Fastcheck build 1227 is still scheduled/pending. Merge protections remain blocked. The earlier encoder-failure note is historical, not the current head's result.

@SolitaryThinker SolitaryThinker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed and verified: the four Wan serve configs parse against current main and resolve to the same FastVideoArgs/SamplingParam values as their source generate configs. Rebased onto main and pushed fixes for the request-field docs, explicit parallelism, README coverage, and a parse test over examples/serving/openai_*.yaml. LGTM.

@SolitaryThinker
SolitaryThinker merged commit 37d06a8 into hao-ai-lab:main Sep 15, 2026
4 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: infra CI, tests, Docker, build type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants