Skip to content

Add a TensorRT-LLM Cosmos3-Edge audiovisual guide - #354

Merged
MaciejBalaNV merged 6 commits into
NVIDIA:mainfrom
ishovkun:docs/cosmos3-edge-trtllm-cookbook
Sep 21, 2026
Merged

MaciejBalaNV merged 6 commits into
NVIDIA:mainfrom
ishovkun:docs/cosmos3-edge-trtllm-cookbook

Conversation

@ishovkun

@ishovkun ishovkun commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add a Cosmos3-Edge section to the TensorRT-LLM audiovisual cookbook, covering exactly what TensorRT-LLM main supports for that checkpoint.

  • add Edge text-to-image, text-to-video, and image-to-video walkthroughs to run_with_trt_llm.ipynb, with the Edge endpoint, launch command, and per-case request shape
  • send Edge's 480p-native contract explicitly: 832x480 with 121 frames, 50 UniPC steps, guidance 5.0 for video, and Edge's native 640x640 with guidance 4.0 for text-to-image. Flow shift (3.0) rides the checkpoint-declared native flow schedule, so requests do not send it
  • run Edge text-to-image in native image mode against /v1/images/generations (see below)
  • document the Edge launch and request contract in the audiovisual README and the shared Cosmos3 environment setup guide, and note the new section in the root notebook index
  • realign the shared audiovisual negative prompt with the TensorRT-LLM default (separate commit, details below)

What TensorRT-LLM head supports for Edge

Checked against main at 8cfa341332. Edge support landed in TensorRT-LLM #16773.

Mode Edge on TensorRT-LLM Basis
Text-to-image supported COSMOS3_EDGE_T2I_PARAMS (640x640, 50 steps, guidance 4.0), via /v1/images/generations with output_type="image"
Text-to-video supported COSMOS3_EDGE_VIDEO_PARAMS (832x480, 121 frames, 50 steps, guidance 5.0, shift 3.0)
Image-to-video supported same video defaults; reference image as multipart input_reference
Synchronized audio not supported Edge has no audio tower
Video-to-video not supported V2V is validated for Nano/Super only
Action (forward/inverse dynamics, policy) not supported the VisualGen pipeline has no action path on main, and the Edge checkpoint's action weights are explicitly unsupported
Transfer controls not supported no transfer/control path exists in TensorRT-LLM main
Reasoning not supported measured: POST /v1/chat/completions returns 404 against a live Edge server, and the served OpenAPI exposes only /health, /metrics, /v1/images/*, /v1/videos/*, /v1/models, /version

Edge is served with no config override — its generation defaults come from the checkpoint. That is also the invocation the upstream test_cosmos3_edge_i2v_example integration test exercises.

For comparison, Edge coverage across the backends already in this repo: Cosmos Framework, Diffusers, and vLLM-Omni ship image-to-video; SGLang ships text-to-image, text-to-video, and image-to-video. This PR brings TensorRT-LLM to the same three modes.

Text-to-image runs in native image mode

Raised in review. The notebook's existing pattern expresses text-to-image as a one-frame video request. At 8cfa341332 that selects video mode, because extra_params.output_type defaults to "video", and the pipeline keys its default negative prompt off that:

def default_negative_prompt(output_type: str) -> str:
    return (COSMOS3_DEFAULT_NEGATIVE_PROMPT if output_type == "image"
            else default_video_negative_prompt())

The example sends no negative prompt of its own, so a still image was being steered by Cosmos3's video negative prompt — "Incoherent motion", "visible frame-to-frame discontinuities", "Movement appears as a slideshow" — none of which mean anything for a single frame.

The Edge case now posts to /v1/images/generations with output_type="image" and decodes the base64 PNG, keeping 640x640, 50 steps, and guidance 4.0. Both halves are required: parse_visual_gen_params does not set output_type for image requests either, so hitting the image endpoint without the flag runs video mode, leaves output.image as None, and returns 500.

For Edge this is the only generation-semantics change — guidance_interval is None in both COSMOS3_EDGE_VIDEO_PARAMS and COSMOS3_EDGE_T2I_PARAMS.

t2i_nano and t2i_super hit the same default and are deliberately left alone here. They are the qwen3 family, whose image-mode table also flips guidance_interval from None to (400, 1000), so switching them changes their output in a way this PR cannot validate (Super needs four GPUs). Worth a follow-up.

The negative-prompt commit

Update after merging main: #353 landed the byte-identical rewording independently, so this change no longer appears in the diff against main. Commit 0407123 remains in the branch history and the explanation below still describes why it was needed at the time.

A prerequisite for the Edge examples to run, but not Edge-specific — it fixes the Nano and Super video cases in the same notebook.

TensorRT-LLM runs the text guardrail over the negative prompt as well as the positive one. The checked-in audiovisual negative prompt had drifted from the reference default in two places — color bleeding between elements and the scene feels lifeless and sterile — and the Cosmos guardrail blocklist matches bleeding and lifeless. Every video request carrying this asset therefore failed with HTTP 500 (Text guardrail blocked prompt) under the documented default of guardrails enabled. Text-to-image was unaffected only because the payload builder attaches no negative prompt for that mode.

TensorRT-LLM #17523 already reworded those exact phrases to color smearing and feels inert in TensorRT-LLM's own default. The cookbook ships its own copy and sends it explicitly as negative_prompt, so the corrected default never applied. This commit adopts the upstream wording; the serialized request body each file produces is now byte-for-byte identical to COSMOS3_VIDEO_NEGATIVE_PROMPT (14674 bytes, key order included).

cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json carries the same drift and is left alone — #310 already covers it.

Validation

Real Edge server, real generations, guardrails enabled throughout.

Environment. 1x NVIDIA B200, x86_64. TensorRT-LLM built from source at 8cfa341332 (reports 1.3.0rc25+8cfa341332). cosmos_guardrail==0.3.0 with the pinned nvidia/Cosmos-1.0-Guardrail revision cf03c0395fac8c4de386c0bdab12cc4fc8d66362. ffmpeg 7.0.2 on PATH, so video responses take the MP4 encoder rather than the AVI fallback. The checkpoint is a local copy of nvidia/Cosmos3-Edge, served by path.

Server. trtllm-serve <Cosmos3-Edge> --port 8000, no --visual_gen_args, exactly as the new section documents.

Execution. The notebook's Edge cells were run verbatim through jupyter nbconvert --to notebook --execute from a checkout of this branch, so find_repo_root, the repo-relative asset paths, create_payload, check_trtllm_server, run_trtllm_payload, and view_run all ran as a reader hits them. nbconvert exited 0 with every cell executed.

Guardrail preflight. All five prompt/negative assets the Edge cases send pass CosmosSafetyChecker.check_text_safety.

Endpoints actually served, from the server log:

1 "POST /v1/images/generations HTTP/1.1" 200 OK     <- t2i_edge
2 "POST /v1/videos/generations HTTP/1.1" 200 OK     <- t2v_edge, i2v_edge

Generation parameters the server resolved, from the server log (warmup line omitted):

Cosmos3 generation dims: 640x640 (WxH), num_frames=1,   num_inference_steps=50, guidance_scale=4.00
Cosmos3 generation dims: 832x480 (WxH), num_frames=121, num_inference_steps=50, guidance_scale=5.00
Cosmos3 generation dims: 832x480 (WxH), num_frames=121, num_inference_steps=50, guidance_scale=5.00

Results, decoded from the returned payloads:

Case Decoded shape FPS Size
t2i_edge (640, 640, 3) PNG n/a 514 KB
t2v_edge (121, 480, 832, 3) 24 510 KB
i2v_edge (121, 480, 832, 3) 24 1952 KB

Every case matches the documented contract. The run logged zero model-card envelope advisories; the one previously emitted came from the one-frame video text-to-image request and disappears in native image mode.

Outputs were inspected visually as well as by shape. As a cross-check the same three cases were also run through cosmos-framework (3845f4c) at matched settings — same prompts, negative prompts, seed, sampler, steps and guidance. Shapes agree on all three, and after the image-mode fix the text-to-image results agree in composition as well, which they did not before: cosmos-framework already sends an empty negative prompt for text-to-image, so it was the TensorRT-LLM video-mode default that diverged.

TensorRT-LLM serves Cosmos3-Edge for text-to-image, text-to-video, and
image-to-video (TensorRT-LLM #16773). Edge has no audio tower, its action
weights are not served by the VisualGen pipeline, and video-to-video is
validated for Nano and Super only, so the new section covers exactly the
three supported modes.

- add a Cosmos3-Edge section to the audiovisual TensorRT-LLM notebook with
  text-to-image, text-to-video, and image-to-video examples
- add Edge's 480p-native request shape (832x480 x 121 frames, 50 UniPC steps,
  guidance 5.0; 640x640 with guidance 4.0 for text-to-image) as explicit
  per-case sampling, and the matching endpoint and launch command
- document the Edge launch and request contract in the audiovisual README and
  the shared Cosmos3 environment setup guide
…sorRT-LLM default

The shared audiovisual negative prompt had drifted from the reference default
in two places: "color bleeding between elements" and "the scene feels
lifeless and sterile". TensorRT-LLM reworded those exact phrases to "color
smearing" and "feels inert" in NVIDIA/TensorRT-LLM#17523, because the Cosmos
guardrail blocklist matches "bleeding" and "lifeless".

The cookbook ships its own copy and sends it explicitly as "negative_prompt",
so TensorRT-LLM's corrected default never applies. TensorRT-LLM runs the text
guardrail over the negative prompt as well as the positive one, so every video
request carrying this asset fails with HTTP 500 ("Text guardrail blocked
prompt") under the documented default of guardrails enabled — for Nano and
Super as much as for Edge.

Adopt the upstream wording so the checked-in asset is byte-identical to
COSMOS3_VIDEO_NEGATIVE_PROMPT in
tensorrt_llm/_torch/visual_gen/models/cosmos3/negative_prompt.py.

Verified: the serialized request body each file produces is byte-for-byte equal
to the TensorRT-LLM default (14674 bytes, key order included), and both pass
CosmosSafetyChecker.check_text_safety under cosmos_guardrail 0.3.0 with the
pinned Cosmos-1.0-Guardrail revision.
@ishovkun
ishovkun marked this pull request as ready for review September 16, 2026 03:42

@ConstBob ConstBob left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. The Edge section matches what TensorRT-LLM actually serves (T2I/T2V/I2V only).

Checked the request contract against TensorRT-LLM COSMOS3_EDGE_VIDEO_PARAMS / COSMOS3_EDGE_T2I_PARAMS at 8cfa341332 (832x480, 121 frames, 50 steps, guidance 5.0; T2I 640x640, guidance 4.0, no flow_shift on the wire). Launch with no --visual_gen_args matches the checkpoint-default path. Edge notebook assets resolve (car_driving_edge.json, car_driving.jpg). Compact audiovisual neg_prompt.json is 14674 bytes and no longer contains the guardrail-blocked “bleeding” / “lifeless” phrasing — that fix applies to Nano/Super as well. Did not re-run generations.

Nits, non-blocking:

  • The root notebook-index row still leads with audio + V2V, then mentions Edge; a short “Edge: T2I/T2V/I2V only” would match the notebook intro.
  • EDGE_T2I_SAMPLING still carries num_frames=121 / fps=24 before the text2image override.

- state the Edge section's scope in the root notebook index row instead of
  leaving it implied after the audio/video-to-video description
- give EDGE_T2I_SAMPLING the one-frame request shape (num_frames=1, fps=8)
  it is silently overridden to, so the table reads the same as the wire

Neither changes behaviour: the serialized request bodies for t2i_edge,
t2v_edge, and i2v_edge are byte-identical before and after.
@ishovkun

Copy link
Copy Markdown
Contributor Author

@ConstBob Thanks for the review!
Both fixed in 1b1c5a1.
Index row now states Edge’s scope directly: “…plus a 480p Cosmos3-Edge section limited to text-to-image, text-to-video, and image-to-video (Edge has no audio tower and no video-to-video).”
EDGE_T2I_SAMPLING now carries num_frames=1 / fps=8, matching the one-frame shape it was overridden to, so the table reads the same as the request.
No behavioural change — the serialized bodies for t2i_edge, t2v_edge, and i2v_edge are byte-identical before and after, so the generations you skipped re-running are still covered by the earlier validation.

The Edge text-to-image example followed the notebook's existing pattern of
expressing text-to-image as a one-frame video request. At 8cfa341332 that
selects video mode, because extra_params.output_type defaults to "video", and
the pipeline's default_negative_prompt() keys off output_type: video modes get
Cosmos3's video negative prompt while image modes get an empty one. The example
sends no negative prompt of its own, so a still image was being steered by
motion and frame-to-frame artifact terms.

Post the Edge case to /v1/images/generations with output_type="image" and
decode the base64 PNG, keeping 640x640, 50 steps, and guidance 4.0. The images
API carries no frame or frame-rate fields, so EDGE_T2I_SAMPLING drops them and
view_run() learns to render stills.

For Edge this is the only generation-semantics change: guidance_interval is
None in both COSMOS3_EDGE_VIDEO_PARAMS and COSMOS3_EDGE_T2I_PARAMS.

t2i_nano and t2i_super hit the same default but are left alone here. Their
image-mode tables also flip guidance_interval from None to (400, 1000), a
behavioural change that needs its own validation (Super needs four GPUs).
@ishovkun

ishovkun commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for this. Flagging that it landed on #310, but it describes t2i_edge at 8cfa341332, which is this PR. Quoting it here so this thread stands alone:

  1. The new t2i_edge example sends a one-frame video request without extra_params.output_type="image". At the referenced TensorRT-LLM revision (8cfa341332), this selects video mode and uses its default negative prompt, which differs from native image mode.
  2. Could you switch this to image mode, using the image endpoint and response decoder, while keeping Edge's 640×640 resolution, 50 steps, and guidance 4.0?
  3. If single-frame video generation is intentional, could you clarify in the guide how this differs from native T2I

Confirmed and fixed in 550fc12.

On (1) — correct, and the negative prompt is what bites. output_type defaults to "video", and the pipeline keys its default off it:

def default_negative_prompt(output_type: str) -> str:
    return (COSMOS3_DEFAULT_NEGATIVE_PROMPT if output_type == "image"
            else default_video_negative_prompt())

The example sends no negative prompt of its own, so a still image was being steered by Cosmos3's video negative prompt — "Incoherent motion", "visible frame-to-frame discontinuities", "Movement appears as a slideshow".

On (2) — done. t2i_edge now posts to /v1/images/generations with output_type="image" and decodes the base64 PNG, keeping 640×640, 50 steps, and guidance 4.0. Both halves are required: parse_visual_gen_params does not set output_type for image requests either, so hitting the image endpoint without the flag runs video mode, leaves output.image as None, and returns 500. EDGE_T2I_SAMPLING drops the frame and frame-rate fields since the images API has no slot for them, and view_run() learns to render stills.

For Edge this is the only generation-semantics change — guidance_interval is None in both COSMOS3_EDGE_VIDEO_PARAMS and COSMOS3_EDGE_T2I_PARAMS.

On (3) — moot now that it runs native image mode, but the guide states the distinction explicitly.

Re-validated on a B200 against a source build of 8cfa341332, guardrails enabled, notebook cells run verbatim through nbconvert (exit 0):

1 "POST /v1/images/generations HTTP/1.1" 200 OK     <- t2i_edge
2 "POST /v1/videos/generations HTTP/1.1" 200 OK     <- t2v_edge, i2v_edge

Cosmos3 generation dims: 640x640 (WxH), num_frames=1,   num_inference_steps=50, guidance_scale=4.00
Cosmos3 generation dims: 832x480 (WxH), num_frames=121, num_inference_steps=50, guidance_scale=5.00
Cosmos3 generation dims: 832x480 (WxH), num_frames=121, num_inference_steps=50, guidance_scale=5.00

Model-card envelope advisories went 1 → 0: the one previously logged came from the one-frame video request and disappears in image mode. Cross-checked against cosmos-framework at matched settings (same prompts, seed, sampler, steps, guidance) — text-to-image now agrees in composition, which it did not before. cosmos-framework already sends an empty negative prompt for text-to-image, so the TensorRT-LLM video-mode default was the sole divergence.

One thing I deliberately did not change: t2i_nano and t2i_super hit the same default. They are the qwen3 family, whose image-mode table also flips guidance_interval from None to (400, 1000), so switching them alters their output in a way this PR cannot validate (Super needs four GPUs). Happy to take that as a follow-up.

@skasetty-Siri skasetty-Siri left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rechecked the Edge fix. Native image mode, the image endpoint, and PNG decoding now address my earlier comment. Assisted offline checks passed. I reviewed the author’s updated GPU validation; GPU inference was not independently rerun. No remaining blocking findings in the Edge changes.

ishovkun added a commit to ishovkun/cosmos that referenced this pull request Sep 18, 2026
run_with_trt_llm.ipynb carries the same broken default this PR already
fixed elsewhere: it defaults TRTLLM_ROOT to "$PWD/TensorRT-LLM" in both
the Nano and Super serve blocks, so the --visual_gen_args config path
does not resolve.

It was also self-contradictory about where to stand: the section opens
with "run the commands below from that checkout" and then says "From the
repository root", and only the latter makes the old default correct.
Settle on the checkout, matching the other two files.

Three lines, no additions, to keep the overlap with NVIDIA#354 small.

Verified from a fresh shell with TRTLLM_ROOT unset: the Nano and Super
config paths both resolve with the checkout as the working directory.

Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
…uide

Two content conflicts, both add/add in adjacent prose:

- README.md: keep the Edge-scoped TensorRT-LLM index row and add NVIDIA#353's new
  distilled 4-step row beneath it.
- cookbooks/cosmos3/README.md: chain the TensorRT-LLM PR references so the
  sentence now cites V2V (#16155), the distilled checkpoints (#16563, #16690),
  and Edge (#16773); place the Cosmos3-Edge launch block after Super and before
  the two distilled students, matching the order the vLLM-Omni section already
  uses.

Everything else auto-merged. Verified on the merged tree:

- run_with_trt_llm.ipynb parses, has 106 cells with unique ids and no syntax
  errors, and the three Edge cases produce byte-identical request bodies to
  the pre-merge branch (t2i via /v1/images/generations with
  output_type="image" at 640x640; t2v and i2v via /v1/videos/generations at
  832x480 x 121 frames).
- NVIDIA#353's negative-prompt rewording is identical to 0407123, and both files
  still serialize byte-for-byte to TensorRT-LLM's COSMOS3_VIDEO_NEGATIVE_PROMPT
  (14674 bytes).

Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>

@MaciejBalaNV MaciejBalaNV left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving after the review from @ConstBob and @skasetty-Siri

@MaciejBalaNV
MaciejBalaNV merged commit 003460e into NVIDIA:main Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants