Add a TensorRT-LLM Cosmos3-Edge audiovisual guide - #354
Conversation
TensorRT-LLM serves Cosmos3-Edge for text-to-image, text-to-video, and image-to-video (TensorRT-LLM #16773). Edge has no audio tower, its action weights are not served by the VisualGen pipeline, and video-to-video is validated for Nano and Super only, so the new section covers exactly the three supported modes. - add a Cosmos3-Edge section to the audiovisual TensorRT-LLM notebook with text-to-image, text-to-video, and image-to-video examples - add Edge's 480p-native request shape (832x480 x 121 frames, 50 UniPC steps, guidance 5.0; 640x640 with guidance 4.0 for text-to-image) as explicit per-case sampling, and the matching endpoint and launch command - document the Edge launch and request contract in the audiovisual README and the shared Cosmos3 environment setup guide
…sorRT-LLM default The shared audiovisual negative prompt had drifted from the reference default in two places: "color bleeding between elements" and "the scene feels lifeless and sterile". TensorRT-LLM reworded those exact phrases to "color smearing" and "feels inert" in NVIDIA/TensorRT-LLM#17523, because the Cosmos guardrail blocklist matches "bleeding" and "lifeless". The cookbook ships its own copy and sends it explicitly as "negative_prompt", so TensorRT-LLM's corrected default never applies. TensorRT-LLM runs the text guardrail over the negative prompt as well as the positive one, so every video request carrying this asset fails with HTTP 500 ("Text guardrail blocked prompt") under the documented default of guardrails enabled — for Nano and Super as much as for Edge. Adopt the upstream wording so the checked-in asset is byte-identical to COSMOS3_VIDEO_NEGATIVE_PROMPT in tensorrt_llm/_torch/visual_gen/models/cosmos3/negative_prompt.py. Verified: the serialized request body each file produces is byte-for-byte equal to the TensorRT-LLM default (14674 bytes, key order included), and both pass CosmosSafetyChecker.check_text_safety under cosmos_guardrail 0.3.0 with the pinned Cosmos-1.0-Guardrail revision.
ConstBob
left a comment
There was a problem hiding this comment.
Thanks. The Edge section matches what TensorRT-LLM actually serves (T2I/T2V/I2V only).
Checked the request contract against TensorRT-LLM COSMOS3_EDGE_VIDEO_PARAMS / COSMOS3_EDGE_T2I_PARAMS at 8cfa341332 (832x480, 121 frames, 50 steps, guidance 5.0; T2I 640x640, guidance 4.0, no flow_shift on the wire). Launch with no --visual_gen_args matches the checkpoint-default path. Edge notebook assets resolve (car_driving_edge.json, car_driving.jpg). Compact audiovisual neg_prompt.json is 14674 bytes and no longer contains the guardrail-blocked “bleeding” / “lifeless” phrasing — that fix applies to Nano/Super as well. Did not re-run generations.
Nits, non-blocking:
- The root notebook-index row still leads with audio + V2V, then mentions Edge; a short “Edge: T2I/T2V/I2V only” would match the notebook intro.
EDGE_T2I_SAMPLINGstill carriesnum_frames=121/fps=24before the text2image override.
- state the Edge section's scope in the root notebook index row instead of leaving it implied after the audio/video-to-video description - give EDGE_T2I_SAMPLING the one-frame request shape (num_frames=1, fps=8) it is silently overridden to, so the table reads the same as the wire Neither changes behaviour: the serialized request bodies for t2i_edge, t2v_edge, and i2v_edge are byte-identical before and after.
|
@ConstBob Thanks for the review! |
The Edge text-to-image example followed the notebook's existing pattern of expressing text-to-image as a one-frame video request. At 8cfa341332 that selects video mode, because extra_params.output_type defaults to "video", and the pipeline's default_negative_prompt() keys off output_type: video modes get Cosmos3's video negative prompt while image modes get an empty one. The example sends no negative prompt of its own, so a still image was being steered by motion and frame-to-frame artifact terms. Post the Edge case to /v1/images/generations with output_type="image" and decode the base64 PNG, keeping 640x640, 50 steps, and guidance 4.0. The images API carries no frame or frame-rate fields, so EDGE_T2I_SAMPLING drops them and view_run() learns to render stills. For Edge this is the only generation-semantics change: guidance_interval is None in both COSMOS3_EDGE_VIDEO_PARAMS and COSMOS3_EDGE_T2I_PARAMS. t2i_nano and t2i_super hit the same default but are left alone here. Their image-mode tables also flip guidance_interval from None to (400, 1000), a behavioural change that needs its own validation (Super needs four GPUs).
|
Thanks for this. Flagging that it landed on #310, but it describes
Confirmed and fixed in 550fc12. On (1) — correct, and the negative prompt is what bites. def default_negative_prompt(output_type: str) -> str:
return (COSMOS3_DEFAULT_NEGATIVE_PROMPT if output_type == "image"
else default_video_negative_prompt())The example sends no negative prompt of its own, so a still image was being steered by Cosmos3's video negative prompt — "Incoherent motion", "visible frame-to-frame discontinuities", "Movement appears as a slideshow". On (2) — done. For Edge this is the only generation-semantics change — On (3) — moot now that it runs native image mode, but the guide states the distinction explicitly. Re-validated on a B200 against a source build of Model-card envelope advisories went 1 → 0: the one previously logged came from the one-frame video request and disappears in image mode. Cross-checked against cosmos-framework at matched settings (same prompts, seed, sampler, steps, guidance) — text-to-image now agrees in composition, which it did not before. cosmos-framework already sends an empty negative prompt for text-to-image, so the TensorRT-LLM video-mode default was the sole divergence. One thing I deliberately did not change: |
skasetty-Siri
left a comment
There was a problem hiding this comment.
Rechecked the Edge fix. Native image mode, the image endpoint, and PNG decoding now address my earlier comment. Assisted offline checks passed. I reviewed the author’s updated GPU validation; GPU inference was not independently rerun. No remaining blocking findings in the Edge changes.
run_with_trt_llm.ipynb carries the same broken default this PR already fixed elsewhere: it defaults TRTLLM_ROOT to "$PWD/TensorRT-LLM" in both the Nano and Super serve blocks, so the --visual_gen_args config path does not resolve. It was also self-contradictory about where to stand: the section opens with "run the commands below from that checkout" and then says "From the repository root", and only the latter makes the old default correct. Settle on the checkout, matching the other two files. Three lines, no additions, to keep the overlap with NVIDIA#354 small. Verified from a fresh shell with TRTLLM_ROOT unset: the Nano and Super config paths both resolve with the checkout as the working directory. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
…uide Two content conflicts, both add/add in adjacent prose: - README.md: keep the Edge-scoped TensorRT-LLM index row and add NVIDIA#353's new distilled 4-step row beneath it. - cookbooks/cosmos3/README.md: chain the TensorRT-LLM PR references so the sentence now cites V2V (#16155), the distilled checkpoints (#16563, #16690), and Edge (#16773); place the Cosmos3-Edge launch block after Super and before the two distilled students, matching the order the vLLM-Omni section already uses. Everything else auto-merged. Verified on the merged tree: - run_with_trt_llm.ipynb parses, has 106 cells with unique ids and no syntax errors, and the three Edge cases produce byte-identical request bodies to the pre-merge branch (t2i via /v1/images/generations with output_type="image" at 640x640; t2v and i2v via /v1/videos/generations at 832x480 x 121 frames). - NVIDIA#353's negative-prompt rewording is identical to 0407123, and both files still serialize byte-for-byte to TensorRT-LLM's COSMOS3_VIDEO_NEGATIVE_PROMPT (14674 bytes). Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
MaciejBalaNV
left a comment
There was a problem hiding this comment.
Approving after the review from @ConstBob and @skasetty-Siri
Summary
Add a Cosmos3-Edge section to the TensorRT-LLM audiovisual cookbook, covering exactly what TensorRT-LLM
mainsupports for that checkpoint.run_with_trt_llm.ipynb, with the Edge endpoint, launch command, and per-case request shape/v1/images/generations(see below)What TensorRT-LLM head supports for Edge
Checked against
mainat8cfa341332. Edge support landed in TensorRT-LLM #16773.COSMOS3_EDGE_T2I_PARAMS(640x640, 50 steps, guidance 4.0), via/v1/images/generationswithoutput_type="image"COSMOS3_EDGE_VIDEO_PARAMS(832x480, 121 frames, 50 steps, guidance 5.0, shift 3.0)input_referencemain, and the Edge checkpoint's action weights are explicitly unsupportedmainPOST /v1/chat/completionsreturns 404 against a live Edge server, and the served OpenAPI exposes only/health,/metrics,/v1/images/*,/v1/videos/*,/v1/models,/versionEdge is served with no config override — its generation defaults come from the checkpoint. That is also the invocation the upstream
test_cosmos3_edge_i2v_exampleintegration test exercises.For comparison, Edge coverage across the backends already in this repo: Cosmos Framework, Diffusers, and vLLM-Omni ship image-to-video; SGLang ships text-to-image, text-to-video, and image-to-video. This PR brings TensorRT-LLM to the same three modes.
Text-to-image runs in native image mode
Raised in review. The notebook's existing pattern expresses text-to-image as a one-frame video request. At
8cfa341332that selects video mode, becauseextra_params.output_typedefaults to"video", and the pipeline keys its default negative prompt off that:The example sends no negative prompt of its own, so a still image was being steered by Cosmos3's video negative prompt — "Incoherent motion", "visible frame-to-frame discontinuities", "Movement appears as a slideshow" — none of which mean anything for a single frame.
The Edge case now posts to
/v1/images/generationswithoutput_type="image"and decodes the base64 PNG, keeping 640x640, 50 steps, and guidance 4.0. Both halves are required:parse_visual_gen_paramsdoes not setoutput_typefor image requests either, so hitting the image endpoint without the flag runs video mode, leavesoutput.imageasNone, and returns 500.For Edge this is the only generation-semantics change —
guidance_intervalisNonein bothCOSMOS3_EDGE_VIDEO_PARAMSandCOSMOS3_EDGE_T2I_PARAMS.t2i_nanoandt2i_superhit the same default and are deliberately left alone here. They are theqwen3family, whose image-mode table also flipsguidance_intervalfromNoneto(400, 1000), so switching them changes their output in a way this PR cannot validate (Super needs four GPUs). Worth a follow-up.The negative-prompt commit
Update after merging
main: #353 landed the byte-identical rewording independently, so this change no longer appears in the diff againstmain. Commit0407123remains in the branch history and the explanation below still describes why it was needed at the time.A prerequisite for the Edge examples to run, but not Edge-specific — it fixes the Nano and Super video cases in the same notebook.
TensorRT-LLM runs the text guardrail over the negative prompt as well as the positive one. The checked-in audiovisual negative prompt had drifted from the reference default in two places —
color bleeding between elementsandthe scene feels lifeless and sterile— and the Cosmos guardrail blocklist matchesbleedingandlifeless. Every video request carrying this asset therefore failed with HTTP 500 (Text guardrail blocked prompt) under the documented default of guardrails enabled. Text-to-image was unaffected only because the payload builder attaches no negative prompt for that mode.TensorRT-LLM #17523 already reworded those exact phrases to
color smearingandfeels inertin TensorRT-LLM's own default. The cookbook ships its own copy and sends it explicitly asnegative_prompt, so the corrected default never applied. This commit adopts the upstream wording; the serialized request body each file produces is now byte-for-byte identical toCOSMOS3_VIDEO_NEGATIVE_PROMPT(14674 bytes, key order included).cookbooks/cosmos3/generator/transfer/assets/negative_prompt.jsoncarries the same drift and is left alone — #310 already covers it.Validation
Real Edge server, real generations, guardrails enabled throughout.
Environment. 1x NVIDIA B200, x86_64. TensorRT-LLM built from source at
8cfa341332(reports1.3.0rc25+8cfa341332).cosmos_guardrail==0.3.0with the pinnednvidia/Cosmos-1.0-Guardrailrevisioncf03c0395fac8c4de386c0bdab12cc4fc8d66362.ffmpeg7.0.2 onPATH, so video responses take the MP4 encoder rather than the AVI fallback. The checkpoint is a local copy ofnvidia/Cosmos3-Edge, served by path.Server.
trtllm-serve <Cosmos3-Edge> --port 8000, no--visual_gen_args, exactly as the new section documents.Execution. The notebook's Edge cells were run verbatim through
jupyter nbconvert --to notebook --executefrom a checkout of this branch, sofind_repo_root, the repo-relative asset paths,create_payload,check_trtllm_server,run_trtllm_payload, andview_runall ran as a reader hits them.nbconvertexited 0 with every cell executed.Guardrail preflight. All five prompt/negative assets the Edge cases send pass
CosmosSafetyChecker.check_text_safety.Endpoints actually served, from the server log:
Generation parameters the server resolved, from the server log (warmup line omitted):
Results, decoded from the returned payloads:
t2i_edge(640, 640, 3)PNGt2v_edge(121, 480, 832, 3)i2v_edge(121, 480, 832, 3)Every case matches the documented contract. The run logged zero model-card envelope advisories; the one previously emitted came from the one-frame video text-to-image request and disappears in native image mode.
Outputs were inspected visually as well as by shape. As a cross-check the same three cases were also run through cosmos-framework (
3845f4c) at matched settings — same prompts, negative prompts, seed, sampler, steps and guidance. Shapes agree on all three, and after the image-mode fix the text-to-image results agree in composition as well, which they did not before: cosmos-framework already sends an empty negative prompt for text-to-image, so it was the TensorRT-LLM video-mode default that diverged.