Skip to content

[feat] Define Dreamverse system prompts for T2VA, FL2VA and Ref2VA - #1855

Open
jayzou3773 wants to merge 1 commit into
hao-ai-lab:mainfrom
jayzou3773:codex/dreamverse-h3-prompt-definitions
Open

jayzou3773 wants to merge 1 commit into
hao-ai-lab:mainfrom
jayzou3773:codex/dreamverse-h3-prompt-definitions

Conversation

@jayzou3773

Copy link
Copy Markdown

Purpose

Define a reviewable, mode-aware system-prompt specification for Dreamverse's MiniMax-H3 prompt-rewriting layer.

Related: #1834 and #1835. This is an independent, documentation-only PR based on main; it does not include the generation-mode implementation commits from #1835. Full three-mode application support remains the scope of that separate PR.

The definitions are a draft proposal, not active runtime configuration. Existing LTX prompts and provider behavior are unchanged. No generation-quality improvement is claimed yet.

Changes

  • Define composable common instructions and self-contained T2VA, FL2VA (including first-frame-only/I2VA semantics), and Ref2VA blocks.
  • Separate the inner H3 prompt format from the outer single-clip, continuation, and rollout JSON contracts.
  • Specify caller-provided media observations, runtime-derived reference labels, effective duration, and locked shot plans for keyframe alignment.
  • Document provider output-budget/truncation risks, existing prompt/caller contract mismatches, and validation gates before activation.
  • Summarize seven Reddit discussions alongside the official guides, explicitly distinguishing anecdotal reports, contradictory feedback, model syntax, and product policy choices.
  • Add an original, unvalidated T2VA prompt illustration and a controlled A/B acceptance plan; link the proposal from the Dreamverse developer guide.

Non-goals

  • No runtime or frontend changes, new WebSocket fields, provider API calls, or GPU jobs.
  • No replacement of existing prompt files and no activation of a new enhancer.
  • No claim that the earlier generation-mode smoke test validates these new prompt definitions.

Test Plan

uv tool run --from pre-commit pre-commit run --files \
  docs/design/dreamverse-h3-system-prompt-v1.md \
  docs/design/dreamverse-h3-prompt-research.md \
  docs/contributing/dreamverse-development.md
git diff --cached --check

Also ran a standard-library document consistency check: balanced fenced blocks, existing local Markdown link targets, four parseable JSON contract examples, and the required ordered section names in each of the three mode definitions.

Test Results

codespell........................................................Passed
PyMarkdown.......................................................Passed
Check for spaces in all filenames................................Passed
Suggestion.......................................................Passed
yapf / ruff / actionlint / mypy: no applicable files (skipped)
Document consistency: 3 documents, local links, 4 JSON contracts,
and 3 self-contained mode field orders — PASS
git diff --cached --check — PASS

GPU inference, prompt A/B evaluation, and a full MkDocs site build were not run for this definition-only change. Integration and generated-media quality acceptance remain explicit follow-up work.

Checklist

  • Ran the project pre-commit hook chain for all changed files.
  • Checked document links, example JSON, and mode-specific field ordering.
  • Updated the developer documentation entry point.
  • Considered GPU memory impact: no runtime changes or allocations in this PR.
  • Activate and validate the new definitions in a separate implementation change.

@mergify mergify Bot added type: docs Documentation only scope: docs Documentation labels Sep 15, 2026
@mergify

mergify Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI

Protection Waiting on
🔴 PR merge requirements 👀 reviews and 🤖 CI

🔴 PR merge requirements

Waiting for

  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • check-success~=pre-commit
This rule is failing.
  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • check-success~=pre-commit
  • title~=(?i)^\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\]

@jayzou3773

Copy link
Copy Markdown
Author

Full-H3 prompt-definition visual examples

Visual examples associated with the mode-aware prompt definitions. These are real, non-mock Full-H3 outputs generated on NVIDIA B200 hardware. This PR remains documentation-only: the definitions are not activated as runtime configuration here, and this comment does not rank visual quality.

  • Runtime commit: beb16d036b1dcf5e87639551cf2d0c6170eb8d59
  • Prompt-definition commit: c929f5125e3f3fb4de7a0dc2f48f9fe8350cb141
  • Model revision: 42ed227ee7df40d41602854ae760620d6eb651fe
  • Settings: seed 1000, 1344×768, 124 frames, 24 fps, 50 inference steps
  • Functional checks: H.264 video + stereo AAC audio, 5.207 s, r_frame_rate=24/1, 124 decoded frames, complete FFmpeg audio/video decode

T2VA — text only

Inputs: text prompt only; no conditioning media.
Prompt arm: raw (reused from the earlier verified single-B200 run 20260919161001).

Exact prompt used
A small red toy car rolls from left to right across a plain white tabletop and stops near the right side. Fixed camera, one continuous shot. Complete silence. No text, people, or music.
t2va.mp4

FL2VA — first and last frames

Inputs: Picture 1 as first_frame at 0.00 s; Picture 2 as last_frame at 5.17 s.
Prompt arm: proposed v1; SHA-256 8bb8ce9482e06d6a939b5ebca87395e77fb8376cae0a1896689589000033b2b3.

Exact prompt used
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 5.17-second mark of the target video.

integrated_multimodal_description: [Shot 1] Preserve the natural appearance and high-angle tabletop composition of the supplied frames. Begin at <Picture 1>: the white/silver robot arm holds the tilted dark and pale-gray-green pouring vessel above the clear glass, with the spoon leaning up-right, the open jar just left of the glass, and the separate black gripper at the lower right. Brown liquid continues to enter the glass, and its level rises visibly and progressively. The arm and vessel make only the small continuous positional adjustment needed to bring the outlet toward the ending position near the glass rim. Keep the glass, spoon, jar, gripper, and their tabletop arrangement recognizable and stable while the liquid level changes. In this one continuous steady high-angle shot, progressively approach <Picture 2> and reach its fuller glass and final robot/vessel pose at the end. Do not add people, captions, or overlays. The entire clip is completely silent.
overall_soundscape: N/A
non_diegetic_music: N/A
fl2va.mp4

Ref2VA — ordered reference image

Inputs: Picture 1 as an ordered reference for identity, appearance, and gym setting; it is not treated as a required opening frame.
Prompt arm: proposed v1; SHA-256 b22a646152d44101d3945cdb0bd52cf58fdf4eaf4fee55057a2c13c737f35164.

Exact prompt used
subject_definitions: <Subject 1> is the adult man from <Picture 1>, used for visible identity and appearance: face, short dark hair, athletic build, bare torso, light-gray shorts, and gray trainers. <Subject 2> is the gym setting from <Picture 1>, including its black racks, overhead bars, black brick wall, gray floor, and existing background equipment.
summary: [reference generation] The same man lowers his raised foot and settles into a relaxed upright stance in the same gym, in one fixed full-body shot. The image supplies appearance and setting, not a required opening frame.
retention_analysis: <Subject 1>: partially_preserved from <Picture 1>; retain identity, proportions, hair, and clothing while changing the raised-leg pose into the requested two-feet-down stance. <Subject 2>: fully_preserved from <Picture 1> as the requested setting relationship; maintain the gym's identity and spatial arrangement, without promising pixel-identical reproduction. No audio references are supplied.
detailed_description: Natural live-action appearance derived from <Picture 1>, with the same indoor gym surfaces, clothing colors, and ordinary room lighting. Preserve the source person's visible identity and proportions without imposing a new stylized treatment. [Shot 1] A fixed full-body view centers <Subject 1> in the aisle of <Subject 2>. Frame his head, hands, and both shoes with floor visible beneath him. Black rack uprights flank the aisle, overhead bars continue toward the back, and the black brick wall remains behind him. Use <Picture 1> for the man's face, short dark hair, athletic build, bare torso, light-gray shorts, and gray lace-up trainers with light soles. Keep these recognizable appearance anchors consistent as his pose changes.

Begin with one knee raised and his arms bent near the front of his body, a readable setup for the requested lowering movement. The reference establishes appearance and setting; it does not require a pixel-matched opening frame. His standing leg supports his weight while the raised foot descends toward the floor. Keep the motion compact and continuous, with the knee gradually extending and the descending shoe remaining connected naturally to the leg. His arms relax as he balances, without introducing a separate gesture or another exercise. When the descending shoe makes contact, allow a quiet, brief shoe-contact sound synchronized to that contact. He then settles into a relaxed upright stance with both feet resting on the floor. Let that completed stance remain readable at the end rather than beginning another action.

Maintain the same face, hair, body proportions, shorts, and shoes throughout the movement. The full-body composition keeps the lowering leg and final foot placement visible, without a close-up, cut, camera pan, or zoom. Preserve the gym's spatial arrangement: the dark uprights remain to either side, the overhead bars stay above the aisle, and the wall, hanging equipment, and existing background details remain behind the man. Their role is to establish the same setting, not to supply a new action or story. Do not introduce another person or place an additional prop in his hands. No captions, new signs, or overlays appear. Keep his performance nonverbal; there is no speech, singing, or voiceover. Quiet room ambience may continue beneath the single foot-lowering action, but there is no music. The ending emphasizes the requested result: the same recognizable man standing comfortably with both feet down in the same gym.
overall_soundscape: Quiet gym room ambience and a brief quiet shoe-contact sound synchronized with the foot reaching the floor; no speech, singing, or voiceover.
non_diegetic_music: N/A
ref2va.mp4

Scope

All three files are nonempty and fully decodable. T2VA is included as cross-mode functional context and used the raw prompt shown above; it does not validate the proposed T2VA V1 definition. FL2VA and Ref2VA used the proposed V1 prompts shown above through a separate pinned functional harness. These examples do not activate this documentation-only PR and are not an automated conditioning-fidelity or quality verdict.

@jayzou3773
jayzou3773 marked this pull request as ready for review September 20, 2026 00:59
@greptile-apps

greptile-apps Bot commented Sep 20, 2026

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

This documentation-only PR appears safe to merge, with no actionable correctness, security, or documentation-integrity issues identified.

Summary

This documentation-only PR proposes mode-aware MiniMax-H3 prompt-rewriting contracts without changing runtime behavior.

  • Defines shared rewriting rules and self-contained T2VA, FL2VA, and Ref2VA prompt formats.
  • Separates inner H3 prompt formatting from operation-specific JSON response contracts.
  • Documents required caller context, validation gates, provider budget risks, and a controlled evaluation plan.
  • Adds research notes that distinguish official guidance, anecdotal reports, and Dreamverse product policy.
  • Links the proposal from the existing Dreamverse developer guide.

Greptile automatically discovered a related ticket that helped explain the purpose of this PR: Dreamverse is intended to support the T2VA, FL2VA, and Ref2VA task families while preserving mode-specific conditioning and runtime validation.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[User request] --> C[Common rewriting instructions]
  B[Validated caller context] --> C
  C --> M{Selected H3 format}
  M --> T[T2VA definition]
  M --> F[FL2VA or first-frame-only definition]
  M --> R[Ref2VA definition]
  T --> O{Calling operation}
  F --> O
  R --> O
  O --> S[Single clip: prompt]
  O --> N[Continuation: next_prompt]
  O --> W[Rollout: id, label, segment_prompts]
  S --> V[Application validation]
  N --> V
  W --> V
  V --> H[MiniMax-H3 generation]
Loading

Reviews (1) · Last reviewed commit: "[docs] Define mode-aware H3 system promp..."

@jayzou3773 jayzou3773 changed the title [docs] Define Dreamverse system prompts for T2VA, FL2VA and Ref2VA [feat] Define Dreamverse system prompts for T2VA, FL2VA and Ref2VA Sep 20, 2026
@mergify mergify Bot added the type: feat New feature or capability label Sep 20, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: docs Documentation type: docs Documentation only type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant