Skip to content

[feat] Add fastvideo serve support for FastMetal (Wan on Mac/MLX) — 1.3B, 14B, and 5B - #1802

Open
Ishxn20 wants to merge 4 commits into
hao-ai-lab:mainfrom
Ishxn20:feat/wan-mlx-serving
Open

[feat] Add fastvideo serve support for FastMetal (Wan on Mac/MLX) — 1.3B, 14B, and 5B#1802
Ishxn20 wants to merge 4 commits into
hao-ai-lab:mainfrom
Ishxn20:feat/wan-mlx-serving

Conversation

@Ishxn20

@Ishxn20 Ishxn20 commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Summary

  1. The server couldn't run non-Nvidia models at all. Added the ability to plug in a different generator, so MLX (and anything else non-CUDA) can be served through the same system CUDA already uses. This part re-derives the same capability [feat] Add an H3 server cookbook and prompt playground #1798 (H3 MLX serving) already builds — it touches the same 4 files, so whichever PR merges second will need a quick rebase.
  2. Built the actual Mac pipeline for Wan. It didn't exist before — the only Mac version of Wan was a script that ran once and exited, with no way to keep it loaded and serve requests. 1.3B and 14B share one pipeline (same underlying model); 5B is a different, newer architecture, so it gets its own.

Testing

  • 358 automated tests pass — config loading, bad-request rejection, and routing to the right model size.

  • Confirmed the existing Nvidia server path is untouched.

  • Real generation on Apple Silicon (M1 Pro, 16 GB). 1.3B served via
    fastvideo serve, generated end to end on Metal:

    $ curl -s http://127.0.0.1:50001/v1/videos -H 'Content-Type: application/json' \
        -d '{"model":"fastwan21-1.3b-mlx","prompt":"A fox runs through fresh snow.","seconds":1,"size":"256x256"}'
    
    INFO [wan_pipeline.py:301] Wan MLX denoise step 1/3 complete
    INFO [wan_pipeline.py:301] Wan MLX denoise step 2/3 complete
    INFO [wan_pipeline.py:301] Wan MLX denoise step 3/3 complete
    INFO [video_api.py:225] Video video_gen_1adc5e84... completed in 154.52s
    
    $ ffprobe -v error -count_frames -select_streams v:0 \
        -show_entries stream=nb_read_frames,width,height -of default=nw=1 fox.mp4
    width=256
    height=256
    nb_read_frames=17
    
    1s × 16fps = 16 frames, which Wan rejects (needs 1 mod 4); the request
    produced 17. Video attached in the comments.
    
fox.mp4
  • 5B prompt-encoder parity — bit-identical. Compared the server's
    _encode_wan_prompt against the reference script's encode_prompt for the
    same prompt on the 5B recipe (fp16 / CPU):

    server: torch.float16 (1, 512, 4096)
    script: torch.float16 (1, 512, 4096)
    bit-identical: True
    max abs diff: 0.0
    
    Pre-fix the server encoded 5B in bf16 on MPS, losing three mantissa bits
    against the script's fp16.
    
  • 5B (Wan2.2) generation not yet run — needs more unified memory than this
    machine has; a 49-frame 1.3B decode already exhausted 16 GB. Pending a
    larger Mac. (@aryan5v)

Copilot AI lite review requested due to automatic review settings September 1, 2026 03:15
@mergify mergify Bot added scope: inference Inference pipeline, serving, CLI scope: infra CI, tests, Docker, build labels Sep 1, 2026
@mergify

mergify Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

⚠️ PR title format required

Your PR title must start with a type tag in brackets. Examples:

  • [feat] Add new model support
  • [bugfix] Fix VAE tiling corruption
  • [refactor] Restructure training pipeline
  • [perf] Optimize attention kernel
  • [ci] Update test infrastructure
  • [infra] Add activation trace hooks
  • [docs] Add inference guide
  • [misc] Clean up configs
  • [new-model] Port Flux2 to FastVideo
  • [skill] Add add-model agent skill

Valid tags: feat, feature, bugfix, fix, refactor, perf, ci, infra, doc, docs, misc, chore, kernel, new-model, skill, skills

Please update your PR title and the merge protection check will pass automatically.

@mergify

mergify Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI

Protection Waiting on
🔴 PR merge requirements 👀 reviews and 🤖 CI

🔴 PR merge requirements

Waiting for

  • #approved-reviews-by>=1
  • check-success=full-suite-passed
This rule is failing.
  • #approved-reviews-by>=1
  • check-success=full-suite-passed
  • check-success=fastcheck-passed
  • check-success~=pre-commit
  • title~=(?i)^\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\]

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

There are correctness issues in the MLX request validation path (task handling) and the PR also introduces a new serving/runtime surface area that still needs real-hardware validation before it’s safe to merge.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR extends FastVideo’s OpenAI-compatible serving stack to support a non-CUDA runtime (MLX) by allowing the server to plug in an alternate generator implementation and runtime-specific request validation, then adds a native MLX “FastMetal Wan” server/pipeline for 1.3B, 14B, and 5B.

Changes:

  • Add a pluggable generator_factory + per-runtime video_request_validator to the shared OpenAI server app/engine so non-CUDA backends can be served.
  • Introduce MLX Wan pipelines (Wan2.1 for 1.3B/14B and Wan2.2-TI2V for 5B) plus a dedicated MLX Wan server entrypoint.
  • Add tests and example YAML configs covering config parsing, validation, and generator dispatch for all three model sizes.
File summaries
File Description
fastvideo/entrypoints/openai/api_server.py Adds runtime/generator factory hooks and conditionally disables image routes for MLX runtime.
fastvideo/entrypoints/openai/serving_engine.py Generalizes generator shape via ServingGenerator and adds runtime-specific request validation hook.
fastvideo/entrypoints/openai/state.py Updates global generator typing to the new serving generator protocol.
fastvideo/entrypoints/openai/video_api.py Runs runtime-specific request validation before CUDA-oriented model/LoRA validation.
fastvideo/entrypoints/openai/mlx_wan_server.py New MLX Wan server entrypoint, config parsing, request allowlist validation, and MLX generator implementation.
fastvideo/mlx_runtime/wan_pipeline.py New MLX Wan2.1 and Wan2.2-TI2V pipelines and shared prompt/rope helpers for repeated server calls.
fastvideo/tests/entrypoints/test_mlx_wan_server.py Tests MLX Wan server config parsing, allowlist validation, and generator dispatch/routing.
fastvideo/tests/mlx/test_mlx_wan_pipeline.py Tests MLX Wan pipeline constructors’ filesystem + checkpoint-shape validation.
examples/serving/mlx_wan21_1_3b.yaml Example serving config for FastMetal Wan2.1 1.3B.
examples/serving/mlx_wan21_14b.yaml Example serving config for FastMetal Wan2.1 14B.
examples/serving/mlx_wan22_5b.yaml Example serving config for FastMetal Wan2.2-TI2V 5B.
Review details
  • Files reviewed: 11/11 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +89 to +90
if request.task not in (None, "t2v"):
raise ValueError("Wan MLX serving supports task=t2v only.")
Comment thread fastvideo/mlx_runtime/wan_pipeline.py Outdated
Comment on lines +156 to +162
try:
manifest = json.loads(manifest_path.read_text())
except (json.JSONDecodeError, OSError):
return None
config = manifest.get("config", manifest)
channels = config.get("in_channels")
return int(channels) if channels is not None else None
Comment on lines +27 to 30
def get_generator() -> ServingGenerator:
"""Return the global VideoGenerator instance (set during startup)."""
assert _generator is not None, "Server not initialized — generator is None"
return _generator
@SolitaryThinker

Copy link
Copy Markdown
Collaborator

Rebased onto main to pick up #1798’s shared MLX serving infrastructure; no content changes (range-diff clean).

"height",
"fps",
"num_frames",
"seconds",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Align or reject seconds before admitting the job. The shared adapter turns an explicit seconds into num_frames = seconds * fps; with both shipped defaults (16 and 24 fps), that is always 0 mod 4, while plan_refine_resolutions requires Wan frames to be 1 mod 4. A normal OpenAI-style request therefore gets queued and fails only inside generation. Please validate the merged request shape synchronously and either map duration to the nearest valid frame grid (for example seconds * fps + 1) or do not advertise seconds for this runtime.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 11f0c76. validate_wan_video_request now resolves an explicit seconds to a Wan-legal num_frames before the job is admitted, rounding up to the next value that is 1 mod the VAE temporal stride. It mirrors the adapter's explicit-field precedence, including the nested video_params spelling, and leaves an explicit num_frames untouched. create_mlx_wan_app binds the served fps into the validator so alignment uses the config's fps rather than the adapter's generic 24 fallback.

Verified on an M1 Pro: {"seconds": 1, "size": "256x256"} produced 17 frames (ffprobe nb_read_frames=17) where the naive 1×16 = 16 would have been rejected. A {"seconds": 3} request at 480×832 resolved to 49 frames and cleared plan_refine_resolutions plus all three denoise steps; it then hit a Metal OOM in TAEHV decode, which is a memory limit on this 16 GB machine rather than the admission path. Full logs in the comment below.

Comment thread fastvideo/mlx_runtime/wan_pipeline.py Outdated
tokenizer = AutoTokenizer.from_pretrained(model_root / "tokenizer", local_files_only=True)
text_encoder = UMT5EncoderModel.from_pretrained(
model_root / "text_encoder",
torch_dtype=torch.bfloat16,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep the 5B prompt-encoder path on the recipe-validated dtype/device, or prove the new path on hardware. mlx_wan22_generate.py deliberately encodes 5B prompts in FP16 on CPU, but this shared helper forces BF16 on MPS and only casts the already-rounded embeddings back to FP16 afterward. That is not the same math (BF16 loses three mantissa bits), and this PR has no real-Mac generation/parity run to show the 5B output remains valid. Please parameterize dtype/device by model family and retain the 5B FP16 path, with an actual FastMetal-5B smoke/parity result.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 11f0c76. _encode_wan_prompt now takes device_arg and dtype_arg, defaulting to Wan2.1's recipe (bf16 / auto, matching mlx_wan_prompt_to_video.py). MLXWan22Pipeline passes cpu / fp16, so 5B is back on the path mlx_wan22_generate.py validated. The bf16 widening before the NumPy hand-off is now conditional, since fp16 maps directly.

Parity result on an M1 Pro — server path vs. the reference script, same prompt, both on the 5B recipe:
server: torch.float16 (1, 512, 4096)
script: torch.float16 (1, 512, 4096)
bit-identical: True
max abs diff: 0.0

@SolitaryThinker SolitaryThinker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for the two inline blockers and the still-open task-validation mismatch:

  • An explicit seconds request is admitted but the shared adapter produces a frame count that violates the Wan 1-mod-4 temporal grid under both shipped FPS defaults.
  • The 5B server changes the maintained FP16/CPU prompt-encoder path to BF16/MPS without real-hardware parity evidence.
  • validate_wan_video_request accepts task=t2v, but the shared request adapter rejects every non-None task for non-MiniMax models, as the existing inline review notes.

Please also address the malformed-manifest validation comment, run at least one real Apple-Silicon generation for each distinct pipeline (Wan2.1 and Wan2.2), and prefix the PR title with [feat] so merge protection can pass. Changed-file pre-commit is green on the rebased head; the focused pytest collection is not runnable on this Linux host because package import initializes Triton without an active GPU driver.

@Ishxn20 Ishxn20 changed the title Adds fastvideo serve support for FastMetal (Wan on Mac/MLX) — all three sizes: 1.3B, 14B, and 5B. [feat] Add fastvideo serve support for FastMetal (Wan on Mac/MLX) — 1.3B, 14B, and 5B Sep 7, 2026
@mergify mergify Bot added the type: feat New feature or capability label Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: inference Inference pipeline, serving, CLI scope: infra CI, tests, Docker, build type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants