Skip to content

[feat] Add Dreamverse multimodal inputs and H3 routing - #1835

Open
jayzou3773 wants to merge 6 commits into
hao-ai-lab:mainfrom
jayzou3773:codex/dreamverse-generation-modes
Open

jayzou3773 wants to merge 6 commits into
hao-ai-lab:mainfrom
jayzou3773:codex/dreamverse-generation-modes

Conversation

@jayzou3773

@jayzou3773 jayzou3773 commented Sep 9, 2026 •

Copy link
Copy Markdown

Purpose

Dreamverse now collects and validates the inputs for T2VA, FL2VA, and Ref2VA and passes them through the session/worker boundary to the matching H3 pipeline. For example, selecting FL2VA requires a first-frame image and accepts an optional last frame; selecting Ref2VA preserves the user's image/video/audio reference order.

Related: #1834. This remains a draft pending real Full H3 GPU generation and maintainer review of the Asset List interface.

Changes

  • Add a reusable Asset List, typed initialization payloads, first/last-frame selection, reference ordering, runtime capabilities, project metadata, and actionable validation/recovery.
  • Upload bounded media through HTTP and pass opaque asset IDs over WebSocket. Validate content, size, duration, reference counts, image aspect ratio, and mono/stereo audio. Pin assets until sessions and outstanding worker commands finish.
  • Add the full-h3 profile (full checkpoint, four GPUs/FSDP, 50 steps), route T2VA/FL2VA to the base pipeline and Ref2VA to the reference transformer, and unload the old executor before switching pipelines.
  • Preserve omitted-field legacy behavior and Preview/LTX T2VA support. First/last frames condition the initial FL2VA segment; later segments continue from the previous final frame. Ref2VA reuses references per clip without claiming temporal continuation.
  • Expose the existing MiniMaxH3Reference type lazily through the public fastvideo.api surface.
  • Add matching mock validation, explicit demo labels, browser tests, and an allocation-aware Slurm launcher.

Validation

  • Backend: 132 passed, 13 GPU-marked deselected, using pytest apps/dreamverse/dreamverse/tests/ -m 'not gpu' -q in a local CPU environment.
  • Frontend: 77 passed, 28 existing skipped; typecheck and production build passed. The new protocol tests are active, not added to the existing skipped suite.
  • Generation-mode browser tests: 4/4 desktop Chromium and 4/4 mobile Chromium passed. The final production frontend plus restarted mock also passed 4/4. Covers all three initialization payloads, uploaded media, reference order, decoded playback, same-WebSocket new-project mode changes, and an 11 MiB upload through the Next proxy.
  • Changed-file pre-commit passed. Full-repository pre-commit was run; it still reports the nine pre-existing fastvideo.VideoGenerator / PipelineConfig mypy export errors in eight files. Other full-repository hooks passed after formatting.
  • Real GPU output is not validated. All 18 Slinky nodes were allocated; the demo job remained pending and was cancelled before starting (zero GPU time). The full H3 checkpoint and ARM64 GB200 environment are prepared for the next allocation. Mock clips are synthetic and are not model-quality or performance evidence.

Screenshot

Ref2VA inputs in the explicitly labeled mock demo:

Ref2VA mock input demo

Remaining acceptance

  • Run real Full H3 T2VA, FL2VA, and Ref2VA generation, confirming video and audio output.
  • Verify real base-to-reference and reference-to-base worker reloads and record peak GPU memory/latency.
  • Confirm the Asset List integration and first-required/last-optional FL2VA behavior with maintainers.
  • Obtain required review and green merge checks before marking ready.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Welcome to FastVideo! Thanks for your first pull request.

How our CI works:

PRs run a three-tier CI system:

  1. Pre-commit — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
  2. Fastcheck — six core GPU lanes run automatically via Buildkite (~10-15 min).
  3. Merge gate — a reviewer adds ready; changed paths select only the relevant integration, training, golden, or SSIM coverage.

Before your PR is reviewed:

  • pre-commit run --all-files passes locally
  • You've added or updated tests for your changes
  • The PR description explains what and why

If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.

Useful links:

@mergify mergify Bot added the type: feat New feature or capability label Sep 9, 2026
@mergify

mergify Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI

Protection Waiting on
🔴 PR merge requirements 👀 reviews and 🤖 CI

🔴 PR merge requirements

Waiting for

  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • check-success~=pre-commit
This rule is failing.
  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • check-success~=pre-commit
  • title~=(?i)^\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\]

cursor Bot pushed a commit to aryan5v/FastVideo that referenced this pull request Sep 11, 2026
Align the UI uplift with hao-ai-lab#1835 by mapping creation
studio mode IDs to canonical t2va/fl2va/ref2va wire values in
session_init_v2 and project_init_v1 payloads.

Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
cursor Bot pushed a commit to aryan5v/FastVideo that referenced this pull request Sep 11, 2026
Use canonical t2va/fl2va/ref2va from hao-ai-lab#1834/hao-ai-lab#1835 instead of a parallel
creation_mode field. Wire model, resolution, duration, and image assets
through session_creation_config while leaving H3 mode routing to upstream.

Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
@jayzou3773 jayzou3773 changed the title [feat] Add Dreamverse generation mode selection [feat] Add Dreamverse multimodal inputs and H3 routing Sep 11, 2026
@mergify mergify Bot added the scope: docs Documentation label Sep 11, 2026
@jayzou3773

Copy link
Copy Markdown
Author

GPU validation has been queued for commit beb16d0 (Slurm job 6019). The PR remains Draft while real GPU validation is pending.

The bounded batch job requests one node with 4 GB200 GPUs and a 2-hour runtime limit. It uses a clean, pinned checkout, the prepared CUDA 13 container, and the pinned full-H3 weights. No hosted prompt-enhancement credentials are required for these direct-prompt functional tests.

Planned sequence: T2VA → FL2VA (first/last images) → Ref2VA (image/video/audio references) → T2VA, checking both H3 pipeline-switch directions. The harness rejects mock mode by default and requires four nonempty, decodable audio/video outputs before reporting GPU validation success. Backend logs, per-mode timings, GPU memory/utilization samples, and result summaries will be retained on the test server. The allocation is released on completion, failure, or timeout.

The harness itself has passed a CPU/mock self-test (explicitly NOT GPU validation). Current real GPU status: PENDING (Resources); no inference results are available yet.

@Davids048
Davids048 self-requested a review September 15, 2026 03:42
@jayzou3773

Copy link
Copy Markdown
Author

Real GPU validation completed

Update to the earlier queued-test comment: Slurm job 6019 completed successfully (COMPLETED, exit 0:0) on 2026-09-14, 05:40:03–05:58:45 UTC. Total allocation time was 18m 42s, and the GPUs were released afterward.

Tested configuration

  • Commit: beb16d036b1dcf5e87639551cf2d0c6170eb8d59 (still the PR head when this result was posted); the runner independently verified a clean, pinned checkout.
  • Hardware: 4 × NVIDIA GB200, sequence parallelism 4.
  • Model: full-H3, checkpoint revision 42ed227ee7df40d41602854ae760620d6eb651fe, with base and reference transformers.
  • Runtime: PyTorch 2.12.0+cu130, CUDA 13.0; startup warmup and Torch compilation disabled for this functional smoke test.
  • Result flags: success: true, gpu_validated: true, mock: false, revision_verified_by_runner: true.

Generation results

The test uploaded conditioning assets and exercised the backend over WebSocket, resetting the project between runs within the same session:

Sequence Inputs Result Client wall time Worker-reported end-to-end time
T2VA Text prompt PASS 68.3 s 66.8 s
FL2VA Text + first/last images PASS 68.1 s 66.5 s
Ref2VA Text + image/video/audio references PASS 404.4 s 106.0 s
T2VA again Text prompt; switch back from Ref2VA PASS 112.1 s 62.4 s

All four outputs are approximately 5.207 seconds, 1344 × 768, with H.264 video and stereo AAC audio at 32 kHz. The harness verified stream completion, nonempty media, audio/video tracks with ffprobe, and successful audio/video decoding with ffmpeg. It rejects mock mode by default.

Backend logs confirm switching from the base transformer to transformer_ref and back. Ref2VA's approximately 298 seconds of additional wall time is primarily executor/model-switch initialization and loading, not pure generation time. These numbers are functional-test observations, not performance benchmarks.

Observed peak GPU memory in 10-second samples was 80,033 MiB on GPU 0 and 77,835 MiB on GPUs 1–3. No CUDA OOM or generation failure was found. The first and final T2VA output SHA-256 hashes also match.

Evidence and remaining scope

The four MP4s (01-t2va.mp4 through 04-t2va.mp4), per-output hashes, summary.json, revision/runtime manifest, backend logs, GPU samples, and Slurm completion records are retained with job 6019 on the test server; they are not attached to this comment.

  • This validates real backend generation, asset handling, AV streaming, and both pipeline-switch directions. It does not replace the required repository CI checks or establish visual/audio quality, conditioning fidelity, or browser-driven end-to-end acceptance.
  • Hosted prompt enhancement and prompt-safety behavior were not tested. A Cerebras tcp_warming GET appears in the logs, so this should not be described as a completely network-isolated run.
  • Non-fatal dependency warnings and a shutdown resource-tracker warning about 28 semaphores remain in the logs; the job and validation both exited successfully.

The PR remains Draft pending the remaining review/acceptance work.

@jayzou3773

Copy link
Copy Markdown
Author

Real Full-H3 generation-mode video evidence

Follow-up functional evidence for the three Dreamverse modes. These are real, non-mock Full-H3 outputs generated on NVIDIA B200 hardware; this comment does not rank visual quality.

  • Runtime commit: beb16d036b1dcf5e87639551cf2d0c6170eb8d59
  • Prompt-definition commit: c929f5125e3f3fb4de7a0dc2f48f9fe8350cb141
  • Model revision: 42ed227ee7df40d41602854ae760620d6eb651fe
  • Settings: seed 1000, 1344×768, 124 frames, 24 fps, 50 inference steps
  • Functional checks: H.264 video + stereo AAC audio, 5.207 s, r_frame_rate=24/1, 124 decoded frames, complete FFmpeg audio/video decode

T2VA — text only

Inputs: text prompt only; no conditioning media.
Prompt arm: raw (reused from the earlier verified single-B200 run 20260919161001).

Exact prompt used
A small red toy car rolls from left to right across a plain white tabletop and stops near the right side. Fixed camera, one continuous shot. Complete silence. No text, people, or music.
t2va.mp4

FL2VA — first and last frames

Inputs: Picture 1 as first_frame at 0.00 s; Picture 2 as last_frame at 5.17 s.
Prompt arm: proposed v1; SHA-256 8bb8ce9482e06d6a939b5ebca87395e77fb8376cae0a1896689589000033b2b3.

Exact prompt used
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 5.17-second mark of the target video.

integrated_multimodal_description: [Shot 1] Preserve the natural appearance and high-angle tabletop composition of the supplied frames. Begin at <Picture 1>: the white/silver robot arm holds the tilted dark and pale-gray-green pouring vessel above the clear glass, with the spoon leaning up-right, the open jar just left of the glass, and the separate black gripper at the lower right. Brown liquid continues to enter the glass, and its level rises visibly and progressively. The arm and vessel make only the small continuous positional adjustment needed to bring the outlet toward the ending position near the glass rim. Keep the glass, spoon, jar, gripper, and their tabletop arrangement recognizable and stable while the liquid level changes. In this one continuous steady high-angle shot, progressively approach <Picture 2> and reach its fuller glass and final robot/vessel pose at the end. Do not add people, captions, or overlays. The entire clip is completely silent.
overall_soundscape: N/A
non_diegetic_music: N/A
fl2va.mp4

Ref2VA — ordered reference image

Inputs: Picture 1 as an ordered reference for identity, appearance, and gym setting; it is not treated as a required opening frame.
Prompt arm: proposed v1; SHA-256 b22a646152d44101d3945cdb0bd52cf58fdf4eaf4fee55057a2c13c737f35164.

Exact prompt used
subject_definitions: <Subject 1> is the adult man from <Picture 1>, used for visible identity and appearance: face, short dark hair, athletic build, bare torso, light-gray shorts, and gray trainers. <Subject 2> is the gym setting from <Picture 1>, including its black racks, overhead bars, black brick wall, gray floor, and existing background equipment.
summary: [reference generation] The same man lowers his raised foot and settles into a relaxed upright stance in the same gym, in one fixed full-body shot. The image supplies appearance and setting, not a required opening frame.
retention_analysis: <Subject 1>: partially_preserved from <Picture 1>; retain identity, proportions, hair, and clothing while changing the raised-leg pose into the requested two-feet-down stance. <Subject 2>: fully_preserved from <Picture 1> as the requested setting relationship; maintain the gym's identity and spatial arrangement, without promising pixel-identical reproduction. No audio references are supplied.
detailed_description: Natural live-action appearance derived from <Picture 1>, with the same indoor gym surfaces, clothing colors, and ordinary room lighting. Preserve the source person's visible identity and proportions without imposing a new stylized treatment. [Shot 1] A fixed full-body view centers <Subject 1> in the aisle of <Subject 2>. Frame his head, hands, and both shoes with floor visible beneath him. Black rack uprights flank the aisle, overhead bars continue toward the back, and the black brick wall remains behind him. Use <Picture 1> for the man's face, short dark hair, athletic build, bare torso, light-gray shorts, and gray lace-up trainers with light soles. Keep these recognizable appearance anchors consistent as his pose changes.

Begin with one knee raised and his arms bent near the front of his body, a readable setup for the requested lowering movement. The reference establishes appearance and setting; it does not require a pixel-matched opening frame. His standing leg supports his weight while the raised foot descends toward the floor. Keep the motion compact and continuous, with the knee gradually extending and the descending shoe remaining connected naturally to the leg. His arms relax as he balances, without introducing a separate gesture or another exercise. When the descending shoe makes contact, allow a quiet, brief shoe-contact sound synchronized to that contact. He then settles into a relaxed upright stance with both feet resting on the floor. Let that completed stance remain readable at the end rather than beginning another action.

Maintain the same face, hair, body proportions, shorts, and shoes throughout the movement. The full-body composition keeps the lowering leg and final foot placement visible, without a close-up, cut, camera pan, or zoom. Preserve the gym's spatial arrangement: the dark uprights remain to either side, the overhead bars stay above the aisle, and the wall, hanging equipment, and existing background details remain behind the man. Their role is to establish the same setting, not to supply a new action or story. Do not introduce another person or place an additional prop in his hands. No captions, new signs, or overlays appear. Keep his performance nonverbal; there is no speech, singing, or voiceover. Quiet room ambience may continue beneath the single foot-lowering action, but there is no music. The ending emphasizes the requested result: the same recognizable man standing comfortably with both feet down in the same gym.
overall_soundscape: Quiet gym room ambience and a brief quiet shoe-contact sound synchronized with the foot reaching the floor; no speech, singing, or voiceover.
non_diegetic_music: N/A
ref2va.mp4

Scope

All three files are nonempty and fully decodable. This is functional generation evidence, not an automated conditioning-fidelity or quality verdict. T2VA used the raw prompt shown above; FL2VA and Ref2VA used the proposed V1 prompts shown above.

@jayzou3773
jayzou3773 marked this pull request as ready for review September 20, 2026 00:59
@greptile-apps

greptile-apps Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 2/5

The PR is not safe to merge until asset access is scoped to an authorized owner/session and failed H3 pipeline switches leave the GPU slot recoverable or unhealthy.

Findings

  1. P1 Security Assets lack access control ▶
  2. P1 Failed switches disable workers ▶
  3. P2 Worker death leaks pins ▶

Summary

This PR adds Dreamverse T2VA, FL2VA, and Ref2VA project inputs, bounded media uploads, typed WebSocket/worker handoff, Full H3 base/reference pipeline routing, frontend asset selection, tests, and Slurm deployment guidance.

  • Adds a runtime-local AssetStore and HTTP upload, preview, and deletion endpoints.
  • Routes T2VA/FL2VA through the H3 base pipeline and Ref2VA through the reference pipeline.
  • Preserves ordered multimodal references and project-scoped asset pinning across the worker boundary.
  • Adds frontend mode selection, Asset List integration, persistence metadata, mock coverage, and browser tests.
  • Greptile automatically discovered a related ticket that helped explain the purpose of this PR: add canonical mode selection, typed conditioning inputs, runtime validation, and matching H3 dispatch while preserving legacy behavior.
Diagram
sequenceDiagram
    participant UI as Dreamverse UI
    participant HTTP as Asset API
    participant WS as Session Controller
    participant Pool as GPU Pool
    participant Worker as GPU Worker
    participant H3 as H3 Pipeline

    UI->>HTTP: POST media bytes
    HTTP-->>UI: asset_id + preview URL
    UI->>WS: session/project init(mode, ordered asset IDs)
    WS->>HTTP: Resolve and pin assets
    WS->>Pool: user_step(GenerationInputs)
    Pool->>Pool: Add step-level pins
    Pool->>Worker: Typed USER_STEP payload
    alt Ref2VA
        Worker->>H3: Load/use reference pipeline
    else T2VA or FL2VA
        Worker->>H3: Load/use base pipeline
    end
    H3-->>Worker: Frames + audio
    Worker-->>Pool: Stream events + StepComplete
    Pool->>Pool: Release step-level pins
    Pool-->>WS: fMP4 stream
    WS-->>UI: Video/audio chunks
Loading

Reviews (1) · Last reviewed commit: "[fix]: open Dreamverse mode menu downwar..."

@router.api_route("/assets/{asset_id}", methods=["GET", "HEAD"])
async def get_asset(asset_id: str) -> FileResponse:
try:
asset = asset_store.get(asset_id)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Assets lack access control

The new GET, HEAD, and DELETE routes resolve assets using only a caller-provided ID from the process-global store. On a shared Dreamverse runtime, a client that obtains another project's asset ID can download that media or delete it while it is unpinned because the store records no owner or session.

How this was verified: Both retrieval and deletion pass the caller-controlled ID directly to the global AssetStore without an identity or ownership check.

Comment on lines 73 to +76
if self.generator is not None:
self.generator.shutdown()
previous_generator = self.generator
self.generator = None
previous_generator.shutdown()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Failed switches disable workers

If loading the base or reference pipeline fails, this code has already cleared and shut down the working generator, but it leaves the pipeline mode and GPU slot health unchanged. Later requests are still assigned to the apparently ready slot, then fail immediately because generate_step rejects the missing generator before it can attempt another pipeline load. The worker remains unusable until it is explicitly reloaded or restarted.

Comment on lines 786 to +788
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
timeout=1800.0)
self._release_step_assets(user_id)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Worker death leaks pins

Step-level asset pins are released only after a worker response. If the worker dies after accepting USER_STEP, the tagged request eventually times out without clearing _step_asset_inputs, and leave_user also skips cleanup when no LeaveAck arrives. Those assets remain undeletable until pool shutdown, and another step for the same user is rejected as still using the previous assets.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: docs Documentation type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant