[feat] Add Dreamverse multimodal inputs and H3 routing - #1835
jayzou3773 wants to merge 6 commits into
Conversation
There was a problem hiding this comment.
Welcome to FastVideo! Thanks for your first pull request.
How our CI works:
PRs run a three-tier CI system:
- Pre-commit — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
- Fastcheck — six core GPU lanes run automatically via Buildkite (~10-15 min).
- Merge gate — a reviewer adds
ready; changed paths select only the relevant integration, training, golden, or SSIM coverage.
Before your PR is reviewed:
-
pre-commit run --all-filespasses locally - You've added or updated tests for your changes
- The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
Useful links:
Merge Protections🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI
🔴 PR merge requirementsWaiting for
This rule is failing.
|
Align the UI uplift with hao-ai-lab#1835 by mapping creation studio mode IDs to canonical t2va/fl2va/ref2va wire values in session_init_v2 and project_init_v1 payloads. Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Use canonical t2va/fl2va/ref2va from hao-ai-lab#1834/hao-ai-lab#1835 instead of a parallel creation_mode field. Wire model, resolution, duration, and image assets through session_creation_config while leaving H3 mode routing to upstream. Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
|
GPU validation has been queued for commit beb16d0 (Slurm job 6019). The PR remains Draft while real GPU validation is pending. The bounded batch job requests one node with 4 GB200 GPUs and a 2-hour runtime limit. It uses a clean, pinned checkout, the prepared CUDA 13 container, and the pinned full-H3 weights. No hosted prompt-enhancement credentials are required for these direct-prompt functional tests. Planned sequence: T2VA → FL2VA (first/last images) → Ref2VA (image/video/audio references) → T2VA, checking both H3 pipeline-switch directions. The harness rejects mock mode by default and requires four nonempty, decodable audio/video outputs before reporting GPU validation success. Backend logs, per-mode timings, GPU memory/utilization samples, and result summaries will be retained on the test server. The allocation is released on completion, failure, or timeout. The harness itself has passed a CPU/mock self-test (explicitly NOT GPU validation). Current real GPU status: PENDING (Resources); no inference results are available yet. |
Real GPU validation completedUpdate to the earlier queued-test comment: Slurm job Tested configuration
Generation resultsThe test uploaded conditioning assets and exercised the backend over WebSocket, resetting the project between runs within the same session:
All four outputs are approximately 5.207 seconds, 1344 × 768, with H.264 video and stereo AAC audio at 32 kHz. The harness verified stream completion, nonempty media, audio/video tracks with Backend logs confirm switching from the base transformer to Observed peak GPU memory in 10-second samples was 80,033 MiB on GPU 0 and 77,835 MiB on GPUs 1–3. No CUDA OOM or generation failure was found. The first and final T2VA output SHA-256 hashes also match. Evidence and remaining scopeThe four MP4s (
The PR remains Draft pending the remaining review/acceptance work. |
Real Full-H3 generation-mode video evidenceFollow-up functional evidence for the three Dreamverse modes. These are real, non-mock Full-H3 outputs generated on NVIDIA B200 hardware; this comment does not rank visual quality.
T2VA — text onlyInputs: text prompt only; no conditioning media. Exact prompt usedt2va.mp4FL2VA — first and last framesInputs: Exact prompt usedfl2va.mp4Ref2VA — ordered reference imageInputs: Exact prompt usedref2va.mp4ScopeAll three files are nonempty and fully decodable. This is functional generation evidence, not an automated conditioning-fidelity or quality verdict. T2VA used the raw prompt shown above; FL2VA and Ref2VA used the proposed V1 prompts shown above. |
|
| @router.api_route("/assets/{asset_id}", methods=["GET", "HEAD"]) | ||
| async def get_asset(asset_id: str) -> FileResponse: | ||
| try: | ||
| asset = asset_store.get(asset_id) |
There was a problem hiding this comment.
The new GET, HEAD, and DELETE routes resolve assets using only a caller-provided ID from the process-global store. On a shared Dreamverse runtime, a client that obtains another project's asset ID can download that media or delete it while it is unpinned because the store records no owner or session.
How this was verified: Both retrieval and deletion pass the caller-controlled ID directly to the global AssetStore without an identity or ownership check.
| if self.generator is not None: | ||
| self.generator.shutdown() | ||
| previous_generator = self.generator | ||
| self.generator = None | ||
| previous_generator.shutdown() |
There was a problem hiding this comment.
Failed switches disable workers
If loading the base or reference pipeline fails, this code has already cleared and shut down the working generator, but it leaves the pipeline mode and GPU slot health unchanged. Later requests are still assigned to the apparently ready slot, then fail immediately because generate_step rejects the missing generator before it can attempt another pipeline load. The worker remains unusable until it is explicitly reloaded or restarted.
| response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id), | ||
| timeout=1800.0) | ||
| self._release_step_assets(user_id) |
There was a problem hiding this comment.
Step-level asset pins are released only after a worker response. If the worker dies after accepting USER_STEP, the tagged request eventually times out without clearing _step_asset_inputs, and leave_user also skips cleanup when no LeaveAck arrives. Those assets remain undeletable until pool shutdown, and another step for the same user is rejected as still using the previous assets.
Purpose
Dreamverse now collects and validates the inputs for T2VA, FL2VA, and Ref2VA and passes them through the session/worker boundary to the matching H3 pipeline. For example, selecting FL2VA requires a first-frame image and accepts an optional last frame; selecting Ref2VA preserves the user's image/video/audio reference order.
Related: #1834. This remains a draft pending real Full H3 GPU generation and maintainer review of the Asset List interface.
Changes
full-h3profile (full checkpoint, four GPUs/FSDP, 50 steps), route T2VA/FL2VA to the base pipeline and Ref2VA to the reference transformer, and unload the old executor before switching pipelines.MiniMaxH3Referencetype lazily through the publicfastvideo.apisurface.Validation
pytest apps/dreamverse/dreamverse/tests/ -m 'not gpu' -qin a local CPU environment.fastvideo.VideoGenerator/PipelineConfigmypy export errors in eight files. Other full-repository hooks passed after formatting.Screenshot
Ref2VA inputs in the explicitly labeled mock demo:
Remaining acceptance