From f5deb1c691be90587a5e346df5a0dd75cb961d1a Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 7 Aug 2026 14:09:17 -0700 Subject: [PATCH 01/14] Add TensorRT-LLM audio and V2V guide Signed-off-by: Igor Shovkun --- README.md | 2 +- cookbooks/cosmos3/README.md | 34 +- .../cosmos3/generator/audiovisual/README.md | 75 ++- .../audiovisual/run_with_trt_llm.ipynb | 457 +++++++++++++++++- 4 files changed, 526 insertions(+), 42 deletions(-) diff --git a/README.md b/README.md index d753c220..0dce9396 100644 --- a/README.md +++ b/README.md @@ -1141,7 +1141,7 @@ We are building examples that show Cosmos 3 Super/Nano/Edge capabilities end to | --- | --- | --- | --- | --- | | Generator (audiovisual) with Diffusers | Generator | Text-to-image, plus text-to-video and image-to-video each with or without synchronized sound, via `Cosmos3OmniPipeline`. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | | Generator (audiovisual) with Cosmos Framework | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | -| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video, and image-to-video against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | +| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | | Generator (audiovisual) with vLLM-Omni | Generator | Text-to-image, text-to-video, image-to-video, and video-to-video, with supported sound modes, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | | Generator (audiovisual) with NIM | Generator | Text2Video and Image2Video only, against the prebuilt `Cosmos3-Generator` NIM; requests use `POST /v1/infer` and decode JSON `b64_video` responses. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | | Generator (audiovisual) with SGLang | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index e6733a9b..b9f714d7 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -8,7 +8,7 @@ backend you want to run and follow that one section. | --- | --- | --- | | [Cosmos Framework](#cosmos-framework) | Native PyTorch inference, launched with `torchrun` | Reasoner, Generator (Audiovisual, Action, **Transfer**) | | [Diffusers](#diffusers) | Direct generation with `Cosmos3OmniPipeline` | Generator (Audiovisual) | -| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video generation) | Generator (Audiovisual) | +| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video/audio generation) | Generator (Audiovisual) | | [TensorRT-LLM Reasoner](#tensorrt-llm-reasoner) | OpenAI-compatible image/video reasoning server | Reasoner | | [Transformers](#transformers) | Hugging Face Transformers inference | Reasoner | | [vLLM](#vllm) | OpenAI-compatible reasoning server (image/video understanding) | Reasoner | @@ -169,9 +169,12 @@ uv pip install --torch-backend=cu130 \ ## TensorRT-LLM Generator OpenAI-compatible **VisualGen** server for Generator audiovisual text-to-image, -text-to-video, and image-to-video examples. Cosmos3 support was added in TensorRT-LLM PR -[#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824); use a -TensorRT-LLM checkout or package that includes that change. +text-to-video, image-to-video, video-to-video, and synchronized audio examples. +Initial Cosmos3 support was added in TensorRT-LLM PR +[#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824), synchronized audio +in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and +video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155). +Use a TensorRT-LLM checkout or package that includes those changes. Install TensorRT-LLM following its upstream documentation. @@ -182,7 +185,7 @@ Cosmos3 VisualGen change before it is available in your installed package or release image. ```bash -apt-get update && apt-get -y install git git-lfs +apt-get update && apt-get -y install ffmpeg git git-lfs git lfs install git clone https://github.com/NVIDIA/TensorRT-LLM.git @@ -244,14 +247,19 @@ torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \ The server exposes `/health`, `/v1/videos/generations`, `/v1/videos`, and `/v1/images/generations`. The audiovisual notebook uses the validated video -generation endpoint for text-to-image, text-to-video, and image-to-video. Cosmos3 -text-to-image is sent as a one-frame video request, matching the TensorRT-LLM -Cosmos3 pipeline; the notebook sends it as `num_frames=1`, `seconds=1`, and -`fps=8` to satisfy the video request schema while preserving a single generated -frame. Requests send Cosmos3 controls through `extra_params`, -so use a TensorRT-LLM build that includes the Cosmos3 VisualGen API schema. -The notebook sets request-level `max_sequence_length=2048` for longer structured -JSON prompts. +generation endpoint for text-to-image, text-to-video, image-to-video, +video-to-video, and synchronized audio. Cosmos3 text-to-image is sent as a +one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; the notebook +sends it as `num_frames=1`, `seconds=1`, and `fps=8` to satisfy the video request +schema while preserving a single generated frame. Image-to-video and +video-to-video upload their reference media as multipart `input_reference`; +TensorRT-LLM classifies the reference by content. Synchronized audio is enabled +with `enable_audio: true` in `extra_params` and is muxed into the output video. +Keep `ffmpeg` on the server `PATH`: without it, TensorRT-LLM falls back to a +video-only AVI encoder and cannot preserve generated audio. Requests send +Cosmos3 controls through `extra_params`, so use a TensorRT-LLM build that includes +the Cosmos3 VisualGen API schema. The notebook sets request-level +`max_sequence_length=4096` for longer structured JSON prompts. ## TensorRT-LLM Reasoner diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index 0cf12ed4..5d3a21a9 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -258,7 +258,7 @@ the [shared environment setup guide](../../README.md#vllm-omni). ### Quickstart Set up the environment and start the server: -[TensorRT-LLM setup](../../README.md#tensorrt-llm). The notebook targets the +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). The notebook targets the OpenAI-compatible VisualGen API served by `trtllm-serve`. Send a text-to-video request with the synchronous video API: @@ -283,7 +283,7 @@ response = requests.post( "num_frames": 189, "num_inference_steps": 35, "guidance_scale": 6.0, - "max_sequence_length": 2048, + "max_sequence_length": 4096, "seed": 0, "extra_params": { "use_resolution_template": False, @@ -299,8 +299,62 @@ Path(f"/tmp/cosmos3_t2v_trtllm{suffix}").write_bytes(response.content) ``` For image-to-video, post multipart form data to the same endpoint with the -reference image under `input_reference`. TensorRT-LLM Cosmos3 audio/action -generation is not covered by this backend section. +reference image under `input_reference`. To generate synchronized audio for a +text-to-video or image-to-video request, add `"enable_audio": True` to +`extra_params`. Keep `ffmpeg` installed in the server environment so TensorRT-LLM +can mux the generated audio into the MP4; its fallback AVI encoder is video-only. + +For video-to-video, upload an MP4 or AVI reference under `input_reference`. +TensorRT-LLM classifies the upload by content and forwards the encoded bytes to +the Cosmos3 workers, which decode the conditioning window with NVDEC: + +```python +import json +from pathlib import Path + +import requests + +source_video = Path("assets/videos/car_driving_plain.mp4").resolve() +v2v_prompt = json.load(open("assets/prompts/image2video/car_driving.json")) +v2v_negative = json.load(open("assets/negative_prompts/image2video/neg_prompt.json")) + +with source_video.open("rb") as video_file: + response = requests.post( + "http://localhost:8000/v1/videos/generations", + data={ + "prompt": json.dumps(v2v_prompt, ensure_ascii=True, separators=(",", ":")), + "negative_prompt": json.dumps(v2v_negative, ensure_ascii=True, separators=(",", ":")), + "size": "1280x720", + "num_frames": "189", + "fps": "24", + "num_inference_steps": "35", + "guidance_scale": "6.0", + "max_sequence_length": "4096", + "seed": "0", + "extra_params": json.dumps( + { + "use_resolution_template": False, + "use_duration_template": False, + "use_system_prompt": True, + "use_guardrails": True, + "condition_video_latent_indexes": [0, 1], + "condition_video_keep": "first", + }, + separators=(",", ":"), + ), + }, + files={"input_reference": (source_video.name, video_file, "video/mp4")}, + headers={"Accept": "video/mp4, video/x-msvideo"}, + ) +response.raise_for_status() +suffix = ".avi" if "x-msvideo" in response.headers.get("content-type", "") else ".mp4" +Path(f"/tmp/cosmos3_v2v_trtllm{suffix}").write_bytes(response.content) +``` + +`condition_video_latent_indexes` identifies clean latent frames in the output; +with `[0, 1]`, TensorRT-LLM consumes the first five pixel frames from the input. +Set `condition_video_keep` to `"last"` to condition on the corresponding tail +window instead. For text-to-image, use the same video generation endpoint with `num_frames=1`, `seconds=1`, and `fps=8`; TensorRT-LLM Cosmos3 returns a one-frame video @@ -309,16 +363,17 @@ derive an eight-frame clip from `seconds * fps`. The TRT-LLM notebook always sends model-specific `extra_params`, so use a TensorRT-LLM release with the Cosmos3 VisualGen API schema. The notebook sets -request-level `max_sequence_length=2048` for longer structured JSON prompts. +request-level `max_sequence_length=4096` for longer structured JSON prompts. ### Notebook walkthrough [`run_with_trt_llm.ipynb`](./run_with_trt_llm.ipynb) is the full tutorial for the -TensorRT-LLM backend: it walks through text-to-image, text-to-video, and -image-to-video requests against an already-running VisualGen server. Server -launch options (Nano and Super, FP8 dynamic quantization, CFG parallelism, -Ulysses, and parallel VAE) live in the -[shared environment setup guide](../../README.md#tensorrt-llm). +TensorRT-LLM backend: it walks through text-to-image, text-to-video and +image-to-video with or without synchronized audio, and video-to-video requests +against an already-running VisualGen server. Server launch options (Nano and +Super, FP8 dynamic quantization, CFG parallelism, Ulysses, and parallel VAE) +live in the +[shared environment setup guide](../../README.md#tensorrt-llm-generator). ## Run with NIM diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index df6260e7..e087e685 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -17,7 +17,7 @@ "\n", "This notebook calls already-running TensorRT-LLM VisualGen servers with direct `curl` requests from Python.\n", "\n", - "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video, and image-to-video generation.\n" + "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video generation.\n" ], "id": "title" }, @@ -27,7 +27,9 @@ "source": [ "## 1. Prerequisites\n", "\n", - "Use a running TensorRT-LLM server with Cosmos3 VisualGen support and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/generations`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`.\n", + "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/generations`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`. Image-to-video and video-to-video send their reference media as multipart `input_reference`.\n", + "\n", + "Install the `ffmpeg` CLI in the server environment. TensorRT-LLM uses it to encode MP4 and mux synchronized audio; without it, the fallback AVI encoder drops audio.\n", "\n", "Generator requires the Guardrail. Request access to the gated [nvidia/Cosmos-1.0-Guardrail](https://huggingface.co/nvidia/Cosmos-1.0-Guardrail) HF repository before running these examples. TensorRT-LLM loads guardrails by default; to disable them, set `TRTLLM_DISABLE_COSMOS3_GUARDRAILS=1` before starting the server or set `use_guardrails` to `False` in the request `extra_params`.\n", "\n", @@ -46,7 +48,7 @@ "source": [ "## 2. Start the Server\n", "\n", - "Run the TensorRT-LLM VisualGen server before running the request cells. The config YAMLs below come from TensorRT-LLM's Cosmos3 support. To build a checkout from source, follow NVIDIA's [Build from Source](https://nvidia.github.io/TensorRT-LLM/installation/build-from-source.html) guide and then run the commands below from that checkout.\n", + "Run the TensorRT-LLM VisualGen server before running the request cells. The config YAMLs below come from TensorRT-LLM's Cosmos3 support. To build a checkout from source, follow NVIDIA's [Build from Source](https://nvidia.github.io/TensorRT-LLM/installation/build-from-source.html) guide and then run the commands below from that checkout. Install `ffmpeg` in that environment before serving (`apt-get update && apt-get install -y ffmpeg` on Ubuntu/Debian images).\n", "\n", "### Cosmos3-Nano\n", "\n", @@ -210,7 +212,13 @@ " for image_path in sorted(image_dir.iterdir()):\n", " if image_path.suffix.lower() in {\".jpg\", \".jpeg\", \".png\", \".webp\", \".bmp\"}:\n", " print(f\" {image_path.name}\")\n", - " display(Image(filename=str(image_path), width=420))\n" + " display(Image(filename=str(image_path), width=420))\n", + "\n", + "videos_dir = assets_dir / \"videos\"\n", + "print(f\"{videos_dir.relative_to(assets_dir)}:\")\n", + "for video_path in sorted(videos_dir.iterdir()):\n", + " if video_path.suffix.lower() in {\".mp4\", \".avi\"}:\n", + " print(f\" {video_path.name} ({video_path.stat().st_size // 1024} KB)\")\n" ], "id": "preview-code" }, @@ -269,7 +277,7 @@ " \"guidance\": 6.0,\n", " \"fps\": 24,\n", " \"num_frames\": 189,\n", - " \"max_sequence_length\": 2048,\n", + " \"max_sequence_length\": 4096,\n", " \"resolution\": \"720\",\n", " \"aspect_ratio\": \"16,9\",\n", " \"seed\": 0,\n", @@ -294,12 +302,37 @@ " \"mode\": \"text2video\",\n", " \"prompt\": \"assets/prompts/text2video/robot_kitchen.json\",\n", " },\n", + " \"t2av_nano\": {\n", + " \"model\": \"Cosmos3-Nano\",\n", + " \"mode\": \"text2video\",\n", + " \"prompt\": \"assets/prompts/text2video/robot_pouring_water_audio.json\",\n", + " \"enable_audio\": True,\n", + " },\n", " \"i2v_nano\": {\n", " \"model\": \"Cosmos3-Nano\",\n", " \"mode\": \"image2video\",\n", " \"prompt\": \"assets/prompts/image2video/humanoid_robot.json\",\n", " \"image\": \"assets/images/image2video/humanoid_robot.jpg\",\n", " },\n", + " \"i2av_nano\": {\n", + " \"model\": \"Cosmos3-Nano\",\n", + " \"mode\": \"image2video\",\n", + " \"prompt\": \"assets/prompts/image2video/coastal_road_audio.json\",\n", + " \"image\": \"assets/images/image2video/coastal_road_audio.jpg\",\n", + " \"enable_audio\": True,\n", + " },\n", + " \"v2v_nano\": {\n", + " \"model\": \"Cosmos3-Nano\",\n", + " \"mode\": \"video2video\",\n", + " \"prompt\": \"assets/prompts/image2video/car_driving.json\",\n", + " \"negative_prompt\": \"assets/negative_prompts/image2video/neg_prompt.json\",\n", + " \"video\": \"assets/videos/car_driving_plain.mp4\",\n", + " \"extra_params\": {\n", + " \"use_system_prompt\": True,\n", + " \"condition_video_latent_indexes\": [0, 1],\n", + " \"condition_video_keep\": \"first\",\n", + " },\n", + " },\n", " \"t2i_super\": {\n", " \"model\": \"Cosmos3-Super\",\n", " \"mode\": \"text2image\",\n", @@ -310,12 +343,37 @@ " \"mode\": \"text2video\",\n", " \"prompt\": \"assets/prompts/text2video/robot_kitchen.json\",\n", " },\n", + " \"t2av_super\": {\n", + " \"model\": \"Cosmos3-Super\",\n", + " \"mode\": \"text2video\",\n", + " \"prompt\": \"assets/prompts/text2video/robot_pouring_water_audio.json\",\n", + " \"enable_audio\": True,\n", + " },\n", " \"i2v_super\": {\n", " \"model\": \"Cosmos3-Super\",\n", " \"mode\": \"image2video\",\n", " \"prompt\": \"assets/prompts/image2video/humanoid_robot.json\",\n", " \"image\": \"assets/images/image2video/humanoid_robot.jpg\",\n", " },\n", + " \"i2av_super\": {\n", + " \"model\": \"Cosmos3-Super\",\n", + " \"mode\": \"image2video\",\n", + " \"prompt\": \"assets/prompts/image2video/coastal_road_audio.json\",\n", + " \"image\": \"assets/images/image2video/coastal_road_audio.jpg\",\n", + " \"enable_audio\": True,\n", + " },\n", + " \"v2v_super\": {\n", + " \"model\": \"Cosmos3-Super\",\n", + " \"mode\": \"video2video\",\n", + " \"prompt\": \"assets/prompts/image2video/car_driving.json\",\n", + " \"negative_prompt\": \"assets/negative_prompts/image2video/neg_prompt.json\",\n", + " \"video\": \"assets/videos/car_driving_plain.mp4\",\n", + " \"extra_params\": {\n", + " \"use_system_prompt\": True,\n", + " \"condition_video_latent_indexes\": [0, 1],\n", + " \"condition_video_keep\": \"first\",\n", + " },\n", + " },\n", "}\n", "\n", "\n", @@ -354,11 +412,14 @@ "\n", " prompt_path = asset_path(spec[\"prompt\"])\n", " payload_path = payload_dir / f\"{use_case}.json\"\n", + " extra_params = dict(TRTLLM_EXTRA_PARAMS)\n", + " extra_params[\"enable_audio\"] = bool(spec.get(\"enable_audio\", False))\n", + " extra_params.update(spec.get(\"extra_params\", {}))\n", " payload = {\n", " \"model_mode\": spec[\"mode\"],\n", " \"name\": use_case,\n", " \"prompt\": compact_json_file(prompt_path),\n", - " \"extra_params\": dict(TRTLLM_EXTRA_PARAMS),\n", + " \"extra_params\": extra_params,\n", " **FIXED_SAMPLING,\n", " }\n", " if spec[\"mode\"] == \"text2image\":\n", @@ -366,11 +427,16 @@ " payload[\"fps\"] = 8\n", " payload[\"seconds\"] = 1\n", " else:\n", - " negative_prompt_path = asset_path(f\"assets/negative_prompts/{spec['mode']}/neg_prompt.json\")\n", + " negative_prompt_path = asset_path(\n", + " spec.get(\"negative_prompt\") or f\"assets/negative_prompts/{spec['mode']}/neg_prompt.json\"\n", + " )\n", " payload[\"negative_prompt\"] = compact_json_file(negative_prompt_path)\n", " if spec[\"mode\"] == \"image2video\":\n", " image_path = asset_path(spec[\"image\"])\n", " payload[\"vision_path\"] = os.path.relpath(image_path, payload_path.parent)\n", + " if spec[\"mode\"] == \"video2video\":\n", + " video_path = asset_path(spec[\"video\"])\n", + " payload[\"video_path\"] = os.path.relpath(video_path, payload_path.parent)\n", "\n", " payload_path.write_text(json.dumps(payload, indent=2) + \"\\n\")\n", "\n", @@ -387,6 +453,12 @@ " image_display_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", " print(f\"image: {image_display_path.relative_to(COSMOS_ROOT)}\")\n", " display(Image(filename=str(image_display_path), width=420))\n", + " if \"video_path\" in payload:\n", + " video_display_path = resolve_payload_path(payload_path, payload[\"video_path\"])\n", + " print(f\"video: {video_display_path.relative_to(COSMOS_ROOT)}\")\n", + " print(\"note: condition_video_latent_indexes=[0, 1] consumes five input pixel frames\")\n", + " if payload[\"extra_params\"][\"enable_audio\"]:\n", + " print(\"audio: enabled; ffmpeg is required on the server to mux it into MP4\")\n", " preview_keys = [\"model_mode\", \"name\", \"num_steps\", \"guidance\", \"fps\", \"num_frames\", \"max_sequence_length\", \"resolution\", \"aspect_ratio\", \"seed\", \"extra_params\"]\n", " if \"seconds\" in payload:\n", " preview_keys.insert(5, \"seconds\")\n", @@ -473,14 +545,19 @@ " ]\n", " cmd += _auth_headers()\n", "\n", - " if payload[\"model_mode\"] == \"image2video\":\n", + " if payload[\"model_mode\"] in {\"image2video\", \"video2video\"}:\n", " form_body = dict(body)\n", " if \"extra_params\" in form_body:\n", " form_body[\"extra_params\"] = json.dumps(form_body[\"extra_params\"], separators=(\",\", \":\"))\n", " for key, value in form_body.items():\n", " cmd += [\"--form-string\", f\"{key}={value}\"]\n", - " image_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", - " cmd += [\"-F\", f\"input_reference=@{image_path}\"]\n", + " if payload[\"model_mode\"] == \"image2video\":\n", + " reference_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", + " reference_type = \"image/jpeg\"\n", + " else:\n", + " reference_path = resolve_payload_path(payload_path, payload[\"video_path\"])\n", + " reference_type = \"video/mp4\"\n", + " cmd += [\"-F\", f\"input_reference=@{reference_path};type={reference_type}\"]\n", " else:\n", " cmd += [\"-H\", \"Content-Type: application/json\"]\n", " cmd += [\"-d\", json.dumps(body, separators=(\",\", \":\"))]\n", @@ -512,6 +589,8 @@ " print(\"output stem:\", output_stem)\n", " if payload[\"model_mode\"] == \"image2video\":\n", " print(\"input image:\", resolve_payload_path(payload_path, payload[\"vision_path\"]))\n", + " if payload[\"model_mode\"] == \"video2video\":\n", + " print(\"input video:\", resolve_payload_path(payload_path, payload[\"video_path\"]))\n", " t0 = time.time()\n", " output_path = post_video(payload_path=payload_path, payload=payload, output_stem=output_stem, model=model)\n", " print(f\"wrote {output_path} in {time.time() - t0:.1f}s\")\n", @@ -628,9 +707,9 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Nano: Text to Video\n", + "## Nano: Text to Video Without Audio\n", "\n", - "Nano text-to-video generation using a structured JSON prompt.\n" + "Nano text-to-video generation using a structured JSON prompt, with audio disabled.\n" ], "id": "nano-t2v" }, @@ -685,9 +764,66 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Nano: Image to Video\n", + "## Nano: Text to Video with Audio\n", "\n", - "Nano image-to-video generation using its paired image asset.\n" + "Nano text-to-video generation with synchronized audio. The payload enables `enable_audio` in TensorRT-LLM's Cosmos3 `extra_params`; `ffmpeg` muxes the generated audio into the MP4 response.\n" + ], + "id": "nano-t2av" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "t2av_nano_payload, t2av_nano_output, t2av_nano_model = create_payload(\"t2av_nano\")\n" + ], + "id": "nano-t2av-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "nano-t2av-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(t2av_nano_model)\n", + "run_trtllm_payload(t2av_nano_payload, t2av_nano_output, model=t2av_nano_model)\n" + ], + "id": "nano-t2av-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "nano-t2av-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(t2av_nano_output)\n" + ], + "id": "nano-t2av-view" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Nano: Image to Video Without Audio\n", + "\n", + "Nano image-to-video generation using its paired image asset, with audio disabled.\n" ], "id": "nano-i2v" }, @@ -738,6 +874,120 @@ ], "id": "nano-i2v-view" }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Nano: Image to Video with Audio\n", + "\n", + "Nano image-to-video generation using its paired coastal-road image and synchronized audio prompt. The reference image is uploaded as multipart `input_reference`, and `enable_audio` is set in `extra_params`.\n" + ], + "id": "nano-i2av" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "i2av_nano_payload, i2av_nano_output, i2av_nano_model = create_payload(\"i2av_nano\")\n" + ], + "id": "nano-i2av-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "nano-i2av-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(i2av_nano_model)\n", + "run_trtllm_payload(i2av_nano_payload, i2av_nano_output, model=i2av_nano_model)\n" + ], + "id": "nano-i2av-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "nano-i2av-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(i2av_nano_output)\n" + ], + "id": "nano-i2av-view" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Nano: Video to Video\n", + "\n", + "Nano video-to-video generation using a checked-in MP4 reference. TensorRT-LLM classifies the multipart `input_reference` by content and forwards the encoded video to each worker for NVDEC decoding. `condition_video_latent_indexes=[0, 1]` pins the first two output latent frames, which consumes the first five pixel frames of the reference.\n" + ], + "id": "nano-v2v" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "v2v_nano_payload, v2v_nano_output, v2v_nano_model = create_payload(\"v2v_nano\")\n" + ], + "id": "nano-v2v-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "nano-v2v-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(v2v_nano_model)\n", + "run_trtllm_payload(v2v_nano_payload, v2v_nano_output, model=v2v_nano_model)\n" + ], + "id": "nano-v2v-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "nano-v2v-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(v2v_nano_output)\n" + ], + "id": "nano-v2v-view" + }, { "cell_type": "markdown", "metadata": {}, @@ -809,9 +1059,9 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Super: Text to Video\n", + "## Super: Text to Video Without Audio\n", "\n", - "Super text-to-video generation using the same structured JSON prompt.\n" + "Super text-to-video generation using the same structured JSON prompt, with audio disabled.\n" ], "id": "super-t2v" }, @@ -866,9 +1116,66 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Super: Image to Video\n", + "## Super: Text to Video with Audio\n", "\n", - "Super image-to-video generation using its paired image asset.\n" + "Super text-to-video generation with synchronized audio using the same `enable_audio` request contract as Nano.\n" + ], + "id": "super-t2av" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "t2av_super_payload, t2av_super_output, t2av_super_model = create_payload(\"t2av_super\")\n" + ], + "id": "super-t2av-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "super-t2av-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(t2av_super_model)\n", + "run_trtllm_payload(t2av_super_payload, t2av_super_output, model=t2av_super_model)\n" + ], + "id": "super-t2av-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "super-t2av-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(t2av_super_output)\n" + ], + "id": "super-t2av-view" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Super: Image to Video Without Audio\n", + "\n", + "Super image-to-video generation using its paired image asset, with audio disabled.\n" ], "id": "super-i2v" }, @@ -918,6 +1225,120 @@ "view_run(i2v_super_output)\n" ], "id": "super-i2v-view" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Super: Image to Video with Audio\n", + "\n", + "Super image-to-video generation using the paired coastal-road image and synchronized audio prompt.\n" + ], + "id": "super-i2av" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "i2av_super_payload, i2av_super_output, i2av_super_model = create_payload(\"i2av_super\")\n" + ], + "id": "super-i2av-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "super-i2av-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(i2av_super_model)\n", + "run_trtllm_payload(i2av_super_payload, i2av_super_output, model=i2av_super_model)\n" + ], + "id": "super-i2av-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "super-i2av-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(i2av_super_output)\n" + ], + "id": "super-i2av-view" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Super: Video to Video\n", + "\n", + "Super video-to-video generation using the same reference-video request shape and conditioning controls as Nano, against the `Cosmos3-Super` endpoint.\n" + ], + "id": "super-v2v" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "v2v_super_payload, v2v_super_output, v2v_super_model = create_payload(\"v2v_super\")\n" + ], + "id": "super-v2v-payload" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Run\n" + ], + "id": "super-v2v-run-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(v2v_super_model)\n", + "run_trtllm_payload(v2v_super_payload, v2v_super_output, model=v2v_super_model)\n" + ], + "id": "super-v2v-run" + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### View Results\n" + ], + "id": "super-v2v-view-title" + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "view_run(v2v_super_output)\n" + ], + "id": "super-v2v-view" } ], "metadata": { From 7a07a159b8a13ba946bde279c49d0dd393af8f7e Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 7 Aug 2026 15:19:24 -0700 Subject: [PATCH 02/14] Add TensorRT-LLM Transfer, Action, and Distilled guides Signed-off-by: Igor Shovkun --- README.md | 6 +- cookbooks/cosmos3/README.md | 43 +- cookbooks/cosmos3/generator/action/README.md | 84 +++- .../action/run_fd_with_trt_llm.ipynb | 285 ++++++++++++ .../action/run_id_with_trt_llm.ipynb | 283 ++++++++++++ .../action/run_policy_with_trt_llm.ipynb | 324 +++++++++++++ .../cosmos3/generator/audiovisual/README.md | 15 +- .../audiovisual/run_with_trt_llm.ipynb | 231 +++++++++- .../cosmos3/generator/transfer/README.md | 90 +++- .../run_video_transfer_with_trt_llm.ipynb | 425 ++++++++++++++++++ 10 files changed, 1765 insertions(+), 21 deletions(-) create mode 100644 cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb create mode 100644 cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb create mode 100644 cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb create mode 100644 cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb diff --git a/README.md b/README.md index 0dce9396..4ee5568e 100644 --- a/README.md +++ b/README.md @@ -1141,17 +1141,21 @@ We are building examples that show Cosmos 3 Super/Nano/Edge capabilities end to | --- | --- | --- | --- | --- | | Generator (audiovisual) with Diffusers | Generator | Text-to-image, plus text-to-video and image-to-video each with or without synchronized sound, via `Cosmos3OmniPipeline`. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | | Generator (audiovisual) with Cosmos Framework | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | -| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | +| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video and image-to-video with or without synchronized audio, video-to-video, and the published four-step T2I/I2V students against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | | Generator (audiovisual) with vLLM-Omni | Generator | Text-to-image, text-to-video, image-to-video, and video-to-video, with supported sound modes, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | | Generator (audiovisual) with NIM | Generator | Text2Video and Image2Video only, against the prebuilt `Cosmos3-Generator` NIM; requests use `POST /v1/infer` and decode JSON `b64_video` responses. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | | Generator (audiovisual) with SGLang | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | | Forward dynamics with Cosmos Framework | Generator | Forward dynamics: action-conditioned future-observation prediction for AV, DROID, UMI, and human hand pose, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb) | +| Forward dynamics with TensorRT-LLM | Generator | Forward dynamics from a checked-in AV image and action trajectory, returning rollout and action tensors through the VisualGen API. | [Notebook](cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb) | | Forward dynamics with vLLM-Omni | Generator | Forward dynamics: action-conditioned future-observation prediction for AV, DROID, UMI, and human hand pose, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/action/run_fd_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_fd_with_vllm_omni.ipynb) | | Forward dynamics with SGLang | Generator | Forward dynamics: action-conditioned future-observation prediction for AV, DROID, and UMI, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/action/run_fd_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_fd_with_sglang.ipynb) | | Inverse dynamics with Cosmos Framework | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb) | +| Inverse dynamics with TensorRT-LLM | Generator | Inverse dynamics from a checked-in AV observation clip, returning reconstructed video and the predicted trajectory in a tensor payload. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb) | | Inverse dynamics with vLLM-Omni | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_vllm_omni.ipynb) | | Inverse dynamics with SGLang | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | +| Action policy with TensorRT-LLM | Generator | DROID policy inference from a concatenated three-camera first frame, returning a rollout and predicted action chunk through the VisualGen API. | [Notebook](cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | | Transfer with Cosmos Framework | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | +| Transfer with TensorRT-LLM | Generator | Edge, blur, depth, segmentation, and WSM transfer using multipart sources, inline precomputed controls, and synchronous encoded-video responses. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | | Transfer with vLLM-Omni | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | | Reasoner with Cosmos Framework | Reasoner | Text and image reasoning: detailed captioning, robot task planning, 2D grounding, describe-anything, and action-trajectory prompts, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb) | | Reasoner with vLLM | Reasoner | Image and video reasoning: captioning, temporal localization, embodied reasoning, common-sense reasoning, 2D grounding, describe-anything, action CoT, driving scenes, physical-plausibility, and situation understanding, against an OpenAI-compatible vLLM server (Cosmos3-Super on 4 GPUs by default; switch to Nano per the cookbook README). | [Notebook](cookbooks/cosmos3/reasoner/run_with_vllm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/reasoner/run_with_vllm.ipynb) | diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index b9f714d7..0bb705dd 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -8,7 +8,7 @@ backend you want to run and follow that one section. | --- | --- | --- | | [Cosmos Framework](#cosmos-framework) | Native PyTorch inference, launched with `torchrun` | Reasoner, Generator (Audiovisual, Action, **Transfer**) | | [Diffusers](#diffusers) | Direct generation with `Cosmos3OmniPipeline` | Generator (Audiovisual) | -| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video/audio generation) | Generator (Audiovisual) | +| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video/audio/action/transfer generation) | Generator (Audiovisual, Action, **Transfer**) | | [TensorRT-LLM Reasoner](#tensorrt-llm-reasoner) | OpenAI-compatible image/video reasoning server | Reasoner | | [Transformers](#transformers) | Hugging Face Transformers inference | Reasoner | | [vLLM](#vllm) | OpenAI-compatible reasoning server (image/video understanding) | Reasoner | @@ -169,7 +169,8 @@ uv pip install --torch-backend=cu130 \ ## TensorRT-LLM Generator OpenAI-compatible **VisualGen** server for Generator audiovisual text-to-image, -text-to-video, image-to-video, video-to-video, and synchronized audio examples. +text-to-video, image-to-video, video-to-video, synchronized audio, Transfer, and +Action examples, including the published four-step T2I and I2V students. Initial Cosmos3 support was added in TensorRT-LLM PR [#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824), synchronized audio in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and @@ -245,6 +246,24 @@ torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \ --port "$COSMOS3_TRTLLM_PORT" ``` +**Four-step distilled T2I** (single GPU; 1024×1024, one-frame warmup): + +```bash +trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \ + --visual_gen_args "$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-t2i-1gpu.yaml" \ + --port "$COSMOS3_TRTLLM_PORT" +``` + +**Four-step distilled I2V** (single GPU; default 1280×720, 189-frame shape): + +```bash +trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ + --port "$COSMOS3_TRTLLM_PORT" +``` + +Action uses the Nano launch above. For DROID policy inference, replace the model +with `nvidia/Cosmos3-Nano-Policy-DROID` and keep the one-GPU Nano config. + The server exposes `/health`, `/v1/videos/generations`, `/v1/videos`, and `/v1/images/generations`. The audiovisual notebook uses the validated video generation endpoint for text-to-image, text-to-video, image-to-video, @@ -261,6 +280,26 @@ Cosmos3 controls through `extra_params`, so use a TensorRT-LLM build that includ the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. +Transfer uses the synchronous `/v1/videos/generations` route. Upload the source +video as multipart `input_reference`; send `edge`, `blur`, `depth`, `seg`, or +`wsm` under `extra_params`. Precomputed controls are base64-encoded media in the +JSON request and are decoded to bytes at the HTTP boundary. TensorRT-LLM uses +`use_guardrails` for its per-request safety switch; `guardrails`, +`control_path`, and other vLLM-Omni-only names are not interchangeable. + +Action requests use the same synchronous route and upload an image or video as +`input_reference`. Because an action trajectory cannot be represented in MP4 or +AVI, `format=auto` resolves to `safetensors`; the payload contains named `video`, +`action`, and `frame_rate` tensors. The asynchronous `/v1/videos` route also +supports this payload: poll `GET /v1/videos/{id}`, then download it from +`GET /v1/videos/{id}/content`. + +The distilled checkpoints own their four-step stochastic schedules and have +classifier-free guidance baked into their weights. Leave `num_inference_steps` +and `guidance_scale` unset; conflicting values are rejected. Also leave +`use_system_prompt` unset for distilled I2V so its checkpoint-declared default +is applied. + ## TensorRT-LLM Reasoner OpenAI-compatible **reasoning** server for image and video understanding. Run diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 51ea8268..b3abecc7 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -18,15 +18,18 @@ The rest of this doc shows how to run these modes on selected embodiments direct - [Run with Diffusers](#run-with-diffusers) - [Quickstart](#quickstart-1) - [Notebook walkthrough](#diffusers-notebook-walkthrough) -- [Run with vLLM-Omni](#run-with-vllm-omni) +- [Run with TensorRT-LLM](#run-with-tensorrt-llm) - [Quickstart](#quickstart-2) + - [Notebook walkthrough](#tensorrt-llm-notebook-walkthrough) +- [Run with vLLM-Omni](#run-with-vllm-omni) + - [Quickstart](#quickstart-3) - [Notebook walkthrough](#vllm-omni-notebook-walkthrough) - [Post-Train for Cosmos3-Nano-Policy-DROID](#post-train-for-cosmos3-nano-policy-droid) ## Overview -All examples are shown across three different inference backends — native -PyTorch (Cosmos Framework), Diffusers, and vLLM-Omni. Every backend uses the sample +Examples are shown across native PyTorch (Cosmos Framework), Diffusers, +TensorRT-LLM, and vLLM-Omni. Every backend uses the sample assets under [`assets/`](./assets) and covers three tasks: Environment setup for all backends is centralized in the shared @@ -36,7 +39,8 @@ links to the section you need. Generator requires the Guardrail. Request access to the gated [nvidia/Cosmos-1.0-Guardrail](https://huggingface.co/nvidia/Cosmos-1.0-Guardrail) HF repository before running these examples. To disable the guardrail, set -`enable_safety_checker=False` (Diffusers), `guardrails: false` (vLLM-Omni +`enable_safety_checker=False` (Diffusers), `use_guardrails: false` +(TensorRT-LLM `extra_params`), `guardrails: false` (vLLM-Omni `extra_params`/`extra_args`), or `--no-guardrails` (Cosmos Framework). ## Action Definition @@ -165,6 +169,78 @@ before running them. The DROID reader resolves its feature layout from the name it is given, so also set `COSMOS3_DROID_ROOT` to a directory named after the DROID release you have available. +## Run with TensorRT-LLM + +### Quickstart + +Set up and launch the VisualGen server with the one-GPU Nano config: +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). Action requests +upload the conditioning image or video as multipart `input_reference` and put +the mode-specific fields under `extra_params`. + +```python +import json +from pathlib import Path + +import requests +from safetensors.torch import load as load_safetensors + +action_root = Path("cookbooks/cosmos3/generator/action") +image_path = action_root / "assets/images/av_0.jpg" +actions = json.loads((action_root / "assets/actions/av_traj_forward.json").read_text()) + +with image_path.open("rb") as image_file: + response = requests.post( + "http://localhost:8000/v1/videos/generations", + data={ + "prompt": "You are an autonomous vehicle planning system.", + "format": "safetensors", + "seed": "0", + "extra_params": json.dumps( + { + "action_mode": "forward_dynamics", + "domain_name": "av", + "action": actions, + "view_point": "ego_view", + "use_guardrails": True, + } + ), + }, + files={"input_reference": (image_path.name, image_file, "image/jpeg")}, + headers={"Accept": "application/octet-stream"}, + ) +response.raise_for_status() +payload = load_safetensors(response.content) +print(payload["video"].shape, payload["action"].shape, payload["frame_rate"].item()) +``` + +The `av` domain preset supplies its 60-step, 9D, 480p/10 fps recipe; other +recognized domains similarly fill omitted action width, chunk, resolution, and +frame rate. Policy and inverse dynamics predict their action tensors, while +forward dynamics carries the conditioned action alongside the rollout. All +Action modes therefore use a tensor output: `format=auto` selects +`safetensors`, and explicit `mp4`/`avi` is rejected rather than dropping the +trajectory. + +The quickstart uses the blocking `POST /v1/videos/generations` route. The +asynchronous `POST /v1/videos` route returns HTTP 202; poll +`GET /v1/videos/{id}` and download the same tensor payload from +`GET /v1/videos/{id}/content`. + +### TensorRT-LLM notebook walkthrough + +- [`run_fd_with_trt_llm.ipynb`](./run_fd_with_trt_llm.ipynb) — forward + dynamics from the checked-in AV image and 60×9 action trajectory. +- [`run_id_with_trt_llm.ipynb`](./run_id_with_trt_llm.ipynb) — inverse + dynamics from the checked-in AV observation clip. +- [`run_policy_with_trt_llm.ipynb`](./run_policy_with_trt_llm.ipynb) — DROID + policy inference with a concatenated three-camera first frame and + `nvidia/Cosmos3-Nano-Policy-DROID`. + +Each notebook decodes TensorRT-LLM's `safetensors` response, writes the rollout +video, and inspects the named `action` tensor. This response contract is not the +top-level action JSON used by vLLM-Omni. + ## Run with vLLM-Omni ### Quickstart diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb new file mode 100644 index 00000000..62158348 --- /dev/null +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -0,0 +1,285 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "fd-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Action Forward Dynamics with TensorRT-LLM\n", + "\n", + "Run the checked-in autonomous-driving example: a first frame plus a 60-step, 9D action\n", + "trajectory conditions a 61-frame rollout. The `av` domain preset supplies the trained 480p,\n", + "10 fps recipe, so the request does not duplicate those mode defaults.\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", + "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install client-side decoding dependencies in this notebook kernel if needed:\n", + "\n", + "```bash\n", + "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", + "```\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-contract", + "metadata": {}, + "source": [ + "## TensorRT-LLM response contract\n", + "\n", + "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", + "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", + "\n", + "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", + "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", + "`GET /v1/videos/{id}/content`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import json\n", + "import mimetypes\n", + "import time\n", + "\n", + "import imageio.v3 as iio\n", + "import requests\n", + "from IPython.display import Video, display\n", + "from safetensors.torch import load as load_safetensors\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/generations\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", + " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", + " reference_path = reference_path.resolve()\n", + " if not reference_path.exists():\n", + " raise FileNotFoundError(reference_path)\n", + " output_path.parent.mkdir(parents=True, exist_ok=True)\n", + " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", + " form = {\n", + " \"prompt\": prompt,\n", + " \"format\": \"safetensors\",\n", + " \"response_format\": \"url\",\n", + " \"seed\": \"0\",\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"application/octet-stream\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with reference_path.open(\"rb\") as reference_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\")\n", + " if \"application/octet-stream\" not in content_type:\n", + " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", + " output_path.write_bytes(response.content)\n", + " payload = load_safetensors(response.content)\n", + " missing = {\"video\", \"action\"} - payload.keys()\n", + " if missing:\n", + " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", + " print(\"saved tensor payload:\", output_path)\n", + " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", + " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", + " return payload\n", + "\n", + "\n", + "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", + " frames = payload[\"video\"].detach().cpu().numpy()\n", + " if frames.ndim == 5 and frames.shape[0] == 1:\n", + " frames = frames[0]\n", + " if frames.ndim != 4 or frames.shape[-1] != 3:\n", + " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", + " fps_tensor = payload.get(\"frame_rate\")\n", + " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", + " iio.imwrite(output_path, frames, fps=fps)\n", + " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", + " display(Video(str(output_path), embed=True))\n", + " return output_path\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-input-title", + "metadata": {}, + "source": [ + "## Prepare the checked-in AV inputs\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-input", + "metadata": {}, + "outputs": [], + "source": [ + "image_path = ACTION_ROOT / \"assets\" / \"images\" / \"av_0.jpg\"\n", + "action_path = ACTION_ROOT / \"assets\" / \"actions\" / \"av_traj_forward.json\"\n", + "actions = json.loads(action_path.read_text())\n", + "assert len(actions) == 60 and all(len(row) == 9 for row in actions)\n", + "\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", + "extra_params = {\n", + " \"action_mode\": \"forward_dynamics\",\n", + " \"domain_name\": \"av\",\n", + " \"action\": actions,\n", + " \"view_point\": \"ego_view\",\n", + " \"use_guardrails\": True,\n", + "}\n", + "print(\"image:\", image_path)\n", + "print(\"action trajectory:\", action_path, len(actions), \"x\", len(actions[0]))\n", + "print(json.dumps({**extra_params, \"action\": \"<60x9 checked-in trajectory>\"}, indent=2))\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-run-title", + "metadata": {}, + "source": [ + "## Run forward dynamics\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "fd_payload = submit_action(\n", + " prompt,\n", + " image_path,\n", + " extra_params,\n", + " OUTPUT_ROOT / \"forward_dynamics_av.safetensors\",\n", + ")\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-view-title", + "metadata": {}, + "source": [ + "## Inspect the rollout and action tensor\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-view", + "metadata": {}, + "outputs": [], + "source": [ + "save_and_view_video(fd_payload, OUTPUT_ROOT / \"forward_dynamics_av.mp4\")\n", + "fd_action = fd_payload[\"action\"].detach().cpu().numpy()\n", + "print(\"first five action rows:\")\n", + "print(fd_action[:5])\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb new file mode 100644 index 00000000..6258ee4d --- /dev/null +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -0,0 +1,283 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "id-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Action Inverse Dynamics with TensorRT-LLM\n", + "\n", + "Recover an autonomous-driving action trajectory from the checked-in observation clip. The\n", + "`av` domain preset selects the 60-step, 9D trajectory shape and the trained 480p/10 fps recipe.\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", + "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install client-side decoding dependencies in this notebook kernel if needed:\n", + "\n", + "```bash\n", + "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", + "```\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-contract", + "metadata": {}, + "source": [ + "## TensorRT-LLM response contract\n", + "\n", + "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", + "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", + "\n", + "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", + "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", + "`GET /v1/videos/{id}/content`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import json\n", + "import mimetypes\n", + "import time\n", + "\n", + "import imageio.v3 as iio\n", + "import requests\n", + "from IPython.display import Video, display\n", + "from safetensors.torch import load as load_safetensors\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/generations\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", + " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", + " reference_path = reference_path.resolve()\n", + " if not reference_path.exists():\n", + " raise FileNotFoundError(reference_path)\n", + " output_path.parent.mkdir(parents=True, exist_ok=True)\n", + " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", + " form = {\n", + " \"prompt\": prompt,\n", + " \"format\": \"safetensors\",\n", + " \"response_format\": \"url\",\n", + " \"seed\": \"0\",\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"application/octet-stream\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with reference_path.open(\"rb\") as reference_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\")\n", + " if \"application/octet-stream\" not in content_type:\n", + " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", + " output_path.write_bytes(response.content)\n", + " payload = load_safetensors(response.content)\n", + " missing = {\"video\", \"action\"} - payload.keys()\n", + " if missing:\n", + " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", + " print(\"saved tensor payload:\", output_path)\n", + " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", + " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", + " return payload\n", + "\n", + "\n", + "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", + " frames = payload[\"video\"].detach().cpu().numpy()\n", + " if frames.ndim == 5 and frames.shape[0] == 1:\n", + " frames = frames[0]\n", + " if frames.ndim != 4 or frames.shape[-1] != 3:\n", + " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", + " fps_tensor = payload.get(\"frame_rate\")\n", + " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", + " iio.imwrite(output_path, frames, fps=fps)\n", + " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", + " display(Video(str(output_path), embed=True))\n", + " return output_path\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-input-title", + "metadata": {}, + "source": [ + "## Prepare the checked-in AV observation\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-input", + "metadata": {}, + "outputs": [], + "source": [ + "video_path = ACTION_ROOT / \"assets\" / \"videos\" / \"av_0.mp4\"\n", + "if not video_path.exists():\n", + " raise FileNotFoundError(video_path)\n", + "\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", + "extra_params = {\n", + " \"action_mode\": \"inverse_dynamics\",\n", + " \"domain_name\": \"av\",\n", + " \"view_point\": \"ego_view\",\n", + " \"use_guardrails\": True,\n", + "}\n", + "print(\"observation:\", video_path)\n", + "print(json.dumps(extra_params, indent=2))\n", + "display(Video(str(video_path), embed=True))\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-run-title", + "metadata": {}, + "source": [ + "## Run inverse dynamics\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "id_payload = submit_action(\n", + " prompt,\n", + " video_path,\n", + " extra_params,\n", + " OUTPUT_ROOT / \"inverse_dynamics_av.safetensors\",\n", + ")\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-view-title", + "metadata": {}, + "source": [ + "## Inspect the reconstructed video and predicted trajectory\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-view", + "metadata": {}, + "outputs": [], + "source": [ + "save_and_view_video(id_payload, OUTPUT_ROOT / \"inverse_dynamics_av.mp4\")\n", + "id_action = id_payload[\"action\"].detach().cpu().numpy()\n", + "print(\"predicted action shape:\", id_action.shape)\n", + "print(\"first five predicted action rows:\")\n", + "print(id_action[:5])\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb new file mode 100644 index 00000000..f0021dc0 --- /dev/null +++ b/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb @@ -0,0 +1,324 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "policy-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Action Policy with TensorRT-LLM\n", + "\n", + "Build the trained DROID concatenated camera view from checked-in clips, then jointly predict a\n", + "16-step, 10D action chunk and a 17-frame rollout with `Cosmos3-Nano-Policy-DROID`.\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", + "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "trtllm-serve nvidia/Cosmos3-Nano-Policy-DROID \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install client-side decoding dependencies in this notebook kernel if needed:\n", + "\n", + "```bash\n", + "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", + "```\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "policy-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-contract", + "metadata": {}, + "source": [ + "## TensorRT-LLM response contract\n", + "\n", + "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", + "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", + "\n", + "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", + "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", + "`GET /v1/videos/{id}/content`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "policy-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import json\n", + "import mimetypes\n", + "import time\n", + "\n", + "import imageio.v3 as iio\n", + "import requests\n", + "from IPython.display import Video, display\n", + "from safetensors.torch import load as load_safetensors\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/generations\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", + " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", + " reference_path = reference_path.resolve()\n", + " if not reference_path.exists():\n", + " raise FileNotFoundError(reference_path)\n", + " output_path.parent.mkdir(parents=True, exist_ok=True)\n", + " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", + " form = {\n", + " \"prompt\": prompt,\n", + " \"format\": \"safetensors\",\n", + " \"response_format\": \"url\",\n", + " \"seed\": \"0\",\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"application/octet-stream\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with reference_path.open(\"rb\") as reference_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\")\n", + " if \"application/octet-stream\" not in content_type:\n", + " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", + " output_path.write_bytes(response.content)\n", + " payload = load_safetensors(response.content)\n", + " missing = {\"video\", \"action\"} - payload.keys()\n", + " if missing:\n", + " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", + " print(\"saved tensor payload:\", output_path)\n", + " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", + " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", + " return payload\n", + "\n", + "\n", + "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", + " frames = payload[\"video\"].detach().cpu().numpy()\n", + " if frames.ndim == 5 and frames.shape[0] == 1:\n", + " frames = frames[0]\n", + " if frames.ndim != 4 or frames.shape[-1] != 3:\n", + " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", + " fps_tensor = payload.get(\"frame_rate\")\n", + " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", + " iio.imwrite(output_path, frames, fps=fps)\n", + " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", + " display(Video(str(output_path), embed=True))\n", + " return output_path\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-input-title", + "metadata": {}, + "source": [ + "## Build the DROID multiview first frame\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "policy-trt-input", + "metadata": {}, + "outputs": [], + "source": [ + "import subprocess\n", + "\n", + "import imageio_ffmpeg\n", + "from IPython.display import Image as NotebookImage\n", + "from PIL import Image, ImageOps\n", + "\n", + "DROID_ROOT = ACTION_ROOT / \"assets\" / \"droid_lerobot_example\"\n", + "camera_paths = {\n", + " \"wrist\": DROID_ROOT / \"videos/observation.image.wrist_image_left/chunk-000/file-000.mp4\",\n", + " \"left\": DROID_ROOT / \"videos/observation.image.exterior_image_1_left/chunk-000/file-000.mp4\",\n", + " \"right\": DROID_ROOT / \"videos/observation.image.exterior_image_2_left/chunk-000/file-000.mp4\",\n", + "}\n", + "for path in camera_paths.values():\n", + " if not path.exists():\n", + " raise FileNotFoundError(path)\n", + "\n", + "frame_dir = OUTPUT_ROOT / \"policy_inputs\"\n", + "frame_dir.mkdir(parents=True, exist_ok=True)\n", + "ffmpeg = imageio_ffmpeg.get_ffmpeg_exe()\n", + "frames = {}\n", + "for name, video_path in camera_paths.items():\n", + " frame_path = frame_dir / f\"{name}.png\"\n", + " subprocess.run(\n", + " [ffmpeg, \"-y\", \"-loglevel\", \"error\", \"-i\", str(video_path), \"-frames:v\", \"1\", str(frame_path)],\n", + " check=True,\n", + " )\n", + " frames[name] = Image.open(frame_path).convert(\"RGB\")\n", + "\n", + "target_w, target_h = 640, 540\n", + "top_h = target_h // 2\n", + "bottom_h = target_h - top_h\n", + "half_w = target_w // 2\n", + "wrist = ImageOps.fit(frames[\"wrist\"], (target_w, top_h), method=Image.Resampling.BICUBIC)\n", + "left = ImageOps.fit(frames[\"left\"], (half_w, bottom_h), method=Image.Resampling.BICUBIC)\n", + "right = ImageOps.fit(frames[\"right\"], (half_w, bottom_h), method=Image.Resampling.BICUBIC)\n", + "conditioning = Image.new(\"RGB\", (target_w, target_h))\n", + "conditioning.paste(wrist, (0, 0))\n", + "conditioning.paste(left, (0, top_h))\n", + "conditioning.paste(right, (half_w, top_h))\n", + "policy_image_path = frame_dir / \"droid_policy_first_frame.png\"\n", + "conditioning.save(policy_image_path)\n", + "display(NotebookImage(filename=str(policy_image_path), width=640))\n", + "\n", + "prompt = os.environ.get(\n", + " \"COSMOS3_POLICY_PROMPT\",\n", + " \"Pick up the object and place it in the target container.\",\n", + ")\n", + "extra_params = {\n", + " \"action_mode\": \"policy\",\n", + " \"domain_name\": \"droid_lerobot\",\n", + " \"view_point\": \"concat_view\",\n", + " \"use_guardrails\": True,\n", + "}\n", + "print(\"prompt:\", prompt)\n", + "print(json.dumps(extra_params, indent=2))\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-run-title", + "metadata": {}, + "source": [ + "## Run policy inference\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "policy-trt-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "policy_payload = submit_action(\n", + " prompt,\n", + " policy_image_path,\n", + " extra_params,\n", + " OUTPUT_ROOT / \"policy_droid.safetensors\",\n", + ")\n" + ] + }, + { + "cell_type": "markdown", + "id": "policy-trt-view-title", + "metadata": {}, + "source": [ + "## Inspect the rollout and predicted action chunk\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "policy-trt-view", + "metadata": {}, + "outputs": [], + "source": [ + "save_and_view_video(policy_payload, OUTPUT_ROOT / \"policy_droid.mp4\")\n", + "policy_action = policy_payload[\"action\"].detach().cpu().numpy()\n", + "print(\"predicted action shape:\", policy_action.shape)\n", + "print(\"first five predicted action rows:\")\n", + "print(policy_action[:5])\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index 5d3a21a9..abae9eaf 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -361,6 +361,14 @@ For text-to-image, use the same video generation endpoint with `num_frames=1`, response for this path. `num_frames` is passed explicitly so the server does not derive an eight-frame clip from `seconds * fps`. +For `nvidia/Cosmos3-Super-Text2Image-4Step`, use the one-GPU +`cosmos3-t2i-1gpu.yaml` config and a 1024×1024 one-frame request. For +`nvidia/Cosmos3-Super-Image2Video-4Step`, use the checkpoint's default +1280×720, 189-frame, 24 fps deployment shape. Both checkpoints own a fixed +four-step schedule with guidance baked into the weights, so omit +`num_inference_steps` and `guidance_scale`. Leave `use_system_prompt` unset for +distilled I2V so TensorRT-LLM applies the checkpoint-declared default. + The TRT-LLM notebook always sends model-specific `extra_params`, so use a TensorRT-LLM release with the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. @@ -370,9 +378,10 @@ request-level `max_sequence_length=4096` for longer structured JSON prompts. [`run_with_trt_llm.ipynb`](./run_with_trt_llm.ipynb) is the full tutorial for the TensorRT-LLM backend: it walks through text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video requests -against an already-running VisualGen server. Server launch options (Nano and -Super, FP8 dynamic quantization, CFG parallelism, Ulysses, and parallel VAE) -live in the +against an already-running VisualGen server. Independent final sections cover +the published four-step T2I and I2V students without overriding their fixed +sampling recipes. Server launch options (Nano, Super, distilled T2I, and +distilled I2V) live in the [shared environment setup guide](../../README.md#tensorrt-llm-generator). ## Run with NIM diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index e087e685..5c5ca042 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -17,7 +17,7 @@ "\n", "This notebook calls already-running TensorRT-LLM VisualGen servers with direct `curl` requests from Python.\n", "\n", - "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video generation.\n" + "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video and image-to-video with or without synchronized audio, video-to-video generation, and the published four-step T2I/I2V students.\n" ], "id": "title" }, @@ -57,7 +57,9 @@ "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD/TensorRT-LLM}\"\n", "\n", - "trtllm-serve nvidia/Cosmos3-Nano --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" --port 8000\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", "```\n", "\n", "### Cosmos3-Super\n", @@ -67,10 +69,32 @@ "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD/TensorRT-LLM}\"\n", "\n", - "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve nvidia/Cosmos3-Super --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" --port 8000\n", + "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \\\n", + " nvidia/Cosmos3-Super \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" \\\n", + " --port 8000\n", "```\n", "\n", - "TensorRT-LLM exposes `/health` when the server is ready. This notebook sends Cosmos3 model-specific controls through `extra_params`, so use a TensorRT-LLM release that includes the Cosmos3 VisualGen API schema.\n" + "TensorRT-LLM exposes `/health` when the server is ready. This notebook sends Cosmos3 model-specific controls through `extra_params`, so use a TensorRT-LLM release that includes the Cosmos3 VisualGen API schema.\n", + "\n", + "\n", + "### Four-step distilled text-to-image\n", + "\n", + "Use the dedicated one-GPU T2I config so warmup compiles the 1024×1024, one-frame shape:\n", + "\n", + "```bash\n", + "trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-t2i-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "### Four-step distilled image-to-video\n", + "\n", + "The published 720p × 189-frame I2V student uses the default one-GPU deployment:\n", + "\n", + "```bash\n", + "trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step --port 8000\n", + "```\n" ], "id": "start-server" }, @@ -110,6 +134,12 @@ "TRTLLM_ENDPOINTS = {\n", " \"Cosmos3-Nano\": os.environ.get(\"COSMOS3_TRTLLM_NANO_BASE_URL\", DEFAULT_TRTLLM_BASE_URL),\n", " \"Cosmos3-Super\": os.environ.get(\"COSMOS3_TRTLLM_SUPER_BASE_URL\", DEFAULT_TRTLLM_BASE_URL),\n", + " \"Cosmos3-Super-Text2Image-4Step\": os.environ.get(\n", + " \"COSMOS3_TRTLLM_DISTILLED_T2I_BASE_URL\", DEFAULT_TRTLLM_BASE_URL\n", + " ),\n", + " \"Cosmos3-Super-Image2Video-4Step\": os.environ.get(\n", + " \"COSMOS3_TRTLLM_DISTILLED_I2V_BASE_URL\", DEFAULT_TRTLLM_BASE_URL\n", + " ),\n", "}\n", "\n", "os.environ[\"COSMOS3_AUDIOVISUAL_OUTPUT_ROOT\"] = str(COSMOS3_AUDIOVISUAL_OUTPUT_ROOT)\n", @@ -393,6 +423,8 @@ " return 720, 1280\n", " if payload.get(\"resolution\") == \"256\" and payload.get(\"aspect_ratio\") == \"16,9\":\n", " return 192, 320\n", + " if payload.get(\"resolution\") == \"1024\" and payload.get(\"aspect_ratio\") == \"1,1\":\n", + " return 1024, 1024\n", " raise ValueError(f\"Unsupported payload resolution/aspect ratio: {payload.get('resolution')} {payload.get('aspect_ratio')}\")\n", "\n", "\n", @@ -493,11 +525,15 @@ " \"seconds\": payload.get(\"seconds\", payload[\"num_frames\"] / payload[\"fps\"]),\n", " \"fps\": payload[\"fps\"],\n", " \"num_frames\": payload[\"num_frames\"],\n", - " \"num_inference_steps\": payload[\"num_steps\"],\n", - " \"guidance_scale\": payload[\"guidance\"],\n", " \"max_sequence_length\": payload[\"max_sequence_length\"],\n", " \"seed\": payload[\"seed\"],\n", " }\n", + " # Distilled checkpoints own their fixed step count and guidance. Omission is\n", + " # intentional: conflicting request values are rejected by TensorRT-LLM.\n", + " if payload.get(\"num_steps\") is not None:\n", + " body[\"num_inference_steps\"] = payload[\"num_steps\"]\n", + " if payload.get(\"guidance\") is not None:\n", + " body[\"guidance_scale\"] = payload[\"guidance\"]\n", " if payload.get(\"negative_prompt\") is not None:\n", " body[\"negative_prompt\"] = payload[\"negative_prompt\"]\n", " body[\"extra_params\"] = payload[\"extra_params\"]\n", @@ -1339,6 +1375,189 @@ "view_run(v2v_super_output)\n" ], "id": "super-v2v-view" + }, + { + "cell_type": "markdown", + "id": "distilled-title", + "metadata": {}, + "source": [ + "# Four-step Distilled Cosmos3-Super Examples\n", + "\n", + "These checkpoints carry a fixed four-step stochastic schedule and classifier-free guidance baked\n", + "into the weights. The request bodies deliberately omit `num_inference_steps` and\n", + "`guidance_scale`; TensorRT-LLM reads both from the checkpoint and rejects conflicting values.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "DISTILLED_ASSET_SETS = {\n", + " \"t2i_distilled\": {\n", + " \"model\": \"Cosmos3-Super-Text2Image-4Step\",\n", + " \"mode\": \"text2image\",\n", + " \"prompt\": \"assets/prompts/text2image/robot_draping.json\",\n", + " \"resolution\": \"1024\",\n", + " \"aspect_ratio\": \"1,1\",\n", + " \"num_frames\": 1,\n", + " \"fps\": 8,\n", + " \"seconds\": 1,\n", + " },\n", + " \"i2v_distilled\": {\n", + " \"model\": \"Cosmos3-Super-Image2Video-4Step\",\n", + " \"mode\": \"image2video\",\n", + " \"prompt\": \"assets/prompts/image2video/humanoid_robot.json\",\n", + " \"negative_prompt\": \"assets/negative_prompts/image2video/neg_prompt.json\",\n", + " \"image\": \"assets/images/image2video/humanoid_robot.jpg\",\n", + " \"resolution\": \"720\",\n", + " \"aspect_ratio\": \"16,9\",\n", + " \"num_frames\": 189,\n", + " \"fps\": 24,\n", + " },\n", + "}\n", + "\n", + "\n", + "def create_distilled_payload(use_case: str) -> tuple[Path, Path, str]:\n", + " spec = DISTILLED_ASSET_SETS[use_case]\n", + " payload_dir = COSMOS3_AUDIOVISUAL_OUTPUT_ROOT / \"trt_llm\" / \"payloads\" / use_case\n", + " output_dir = COSMOS3_AUDIOVISUAL_OUTPUT_ROOT / \"trt_llm\" / use_case\n", + " payload_dir.mkdir(parents=True, exist_ok=True)\n", + " output_dir.mkdir(parents=True, exist_ok=True)\n", + " payload_path = payload_dir / f\"{use_case}.json\"\n", + " payload = {\n", + " \"model_mode\": spec[\"mode\"],\n", + " \"name\": use_case,\n", + " \"prompt\": compact_json_file(asset_path(spec[\"prompt\"])),\n", + " \"resolution\": spec[\"resolution\"],\n", + " \"aspect_ratio\": spec[\"aspect_ratio\"],\n", + " \"num_frames\": spec[\"num_frames\"],\n", + " \"fps\": spec[\"fps\"],\n", + " \"max_sequence_length\": 4096,\n", + " \"seed\": 0,\n", + " \"extra_params\": {\n", + " \"use_resolution_template\": False,\n", + " \"use_duration_template\": False,\n", + " \"use_guardrails\": True,\n", + " \"enable_audio\": False,\n", + " },\n", + " }\n", + " if \"seconds\" in spec:\n", + " payload[\"seconds\"] = spec[\"seconds\"]\n", + " if \"negative_prompt\" in spec:\n", + " payload[\"negative_prompt\"] = compact_json_file(asset_path(spec[\"negative_prompt\"]))\n", + " if \"image\" in spec:\n", + " image_path = asset_path(spec[\"image\"])\n", + " payload[\"vision_path\"] = os.path.relpath(image_path, payload_path.parent)\n", + " display(Image(filename=str(image_path), width=420))\n", + " # Do not add use_system_prompt: the I2V checkpoint declares True and the\n", + " # pipeline applies that checkpoint default only when the request leaves it unset.\n", + " payload_path.write_text(json.dumps(payload, indent=2) + \"\\n\")\n", + " print(\"model:\", spec[\"model\"])\n", + " print(\"payload:\", payload_path)\n", + " print(\"output:\", output_dir)\n", + " print(json.dumps(payload, indent=2))\n", + " return payload_path, output_dir, spec[\"model\"]\n" + ] + }, + { + "cell_type": "markdown", + "id": "distilled-t2i-title", + "metadata": {}, + "source": [ + "## Distilled text to image\n", + "\n", + "The server is compiled for 1024×1024. The stable video route returns a one-frame media response,\n", + "so the request sends `num_frames=1`, `seconds=1`, and `fps=8` while leaving the fixed sampling\n", + "fields unset.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-t2i-payload", + "metadata": {}, + "outputs": [], + "source": [ + "distilled_t2i_payload, distilled_t2i_output, distilled_t2i_model = create_distilled_payload(\"t2i_distilled\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-t2i-run", + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(distilled_t2i_model)\n", + "distilled_t2i_data = json.loads(distilled_t2i_payload.read_text())\n", + "post_video(\n", + " payload_path=distilled_t2i_payload,\n", + " payload=distilled_t2i_data,\n", + " output_stem=distilled_t2i_output / \"distilled_t2i\",\n", + " model=distilled_t2i_model,\n", + ")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-t2i-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_run(distilled_t2i_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "distilled-i2v-title", + "metadata": {}, + "source": [ + "## Distilled image to video\n", + "\n", + "The published deployment shape is 1280×720 with 189 frames at 24 fps. The request leaves\n", + "`use_system_prompt` unset so the checkpoint's declared `true` default is applied.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-i2v-payload", + "metadata": {}, + "outputs": [], + "source": [ + "distilled_i2v_payload, distilled_i2v_output, distilled_i2v_model = create_distilled_payload(\"i2v_distilled\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-i2v-run", + "metadata": {}, + "outputs": [], + "source": [ + "check_trtllm_server(distilled_i2v_model)\n", + "distilled_i2v_data = json.loads(distilled_i2v_payload.read_text())\n", + "post_video(\n", + " payload_path=distilled_i2v_payload,\n", + " payload=distilled_i2v_data,\n", + " output_stem=distilled_i2v_output / \"distilled_i2v\",\n", + " model=distilled_i2v_model,\n", + ")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "distilled-i2v-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_run(distilled_i2v_output)\n" + ] } ], "metadata": { diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index e3467e4d..f5778a9e 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -2,7 +2,7 @@ Cosmos3 video **transfer** examples — **Nano** (single GPU) and **Super** (multi-GPU, 32B) — on the native PyTorch (Cosmos Framework) path, the Diffusers modular pipeline, and the -OpenAI-compatible vLLM-Omni server path. +OpenAI-compatible TensorRT-LLM and vLLM-Omni server paths. Sample assets under [`assets/`](./assets) cover spatial control signals paired with `prompt.json` files: @@ -13,10 +13,10 @@ Sample assets under [`assets/`](./assets) cover spatial control signals paired w - **World scenario (WSM)** — world-scenario map control plus caption. - **Multi-control** — two or more hints; Cosmos Framework also supports per-hint weights. -Both Cosmos Framework and vLLM-Omni support multi-control transfer. Per-hint -weighting is supported only by Cosmos Framework; vLLM-Omni accepts multiple -controls but does not support per-hint weights. Diffusers accepts multiple -precomputed controls, also without per-hint weights. +Cosmos Framework, TensorRT-LLM, and vLLM-Omni support multi-control transfer. +Per-hint weighting is supported only by Cosmos Framework; the serving backends +accept multiple unweighted controls. Diffusers also accepts multiple +precomputed controls without per-hint weights. Environment setup is centralized in the shared [Cosmos3 cookbooks environment setup](../../README.md) guide. @@ -25,6 +25,8 @@ Environment setup is centralized in the shared Video transfer generates a target clip from a `prompt.json` caption and one or more spatial control signals. The Framework path uses `model_mode` `video2video` in a local JSON spec. +TensorRT-LLM uses `POST /v1/videos/generations`: source media is multipart +`input_reference`, while hints are carried under `extra_params`. The vLLM-Omni path uses `POST /v1/videos/sync` and passes one or more hint keys (`edge`, `blur`, `depth`, `seg`, or `wsm`) inside `extra_params`. Cosmos Framework accepts pre-computed control videos (`control_path`) or derives active controls from a raw source video @@ -191,6 +193,81 @@ by the same five controls on Cosmos3-Super. It reads the same [`specs/`](./specs reuses the previews from [`preview_helpers.py`](./preview_helpers.py), writing outputs to `outputs/notebooks/diffusers//transfer_/vision.mp4`. +## Run with TensorRT-LLM + +### Quickstart + +Set up and launch Nano or Super as described in the shared +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). TensorRT-LLM +accepts source media as multipart `input_reference`. JSON cannot carry bytes, +so a precomputed control is base64 encoded inside its hint object; the server +decodes it at the HTTP boundary before validating and dispatching the request. + +```python +import base64 +import json +from pathlib import Path + +import requests + +transfer_root = Path("cookbooks/cosmos3/generator/transfer") +control_path = transfer_root / "assets/depth/control_depth.mp4" +prompt = json.dumps(json.load(open(transfer_root / "assets/depth/prompt.json"))) +negative = json.dumps(json.load(open(transfer_root / "assets/negative_prompt.json"))) +extra_params = { + "use_resolution_template": False, + "use_duration_template": False, + "use_system_prompt": False, + "use_guardrails": True, + "depth": { + "control": base64.b64encode(control_path.read_bytes()).decode("ascii") + }, + "control_guidance": 1.5, + "num_video_frames_per_chunk": 121, + "num_conditional_frames": 1, + "num_first_chunk_conditional_frames": 0, + "max_frames": 121, +} + +with control_path.open("rb") as source_file: + response = requests.post( + "http://localhost:8000/v1/videos/generations", + data={ + "prompt": prompt, + "negative_prompt": negative, + "size": "1280x720", + "num_frames": "121", + "fps": "30", + "num_inference_steps": "35", + "guidance_scale": "3.0", + "max_sequence_length": "4096", + "seed": "2026", + "extra_params": json.dumps(extra_params), + }, + files={"input_reference": (control_path.name, source_file, "video/mp4")}, + headers={"Accept": "video/mp4, video/x-msvideo"}, + ) +response.raise_for_status() +suffix = ".avi" if "avi" in response.headers.get("content-type", "") else ".mp4" +Path(f"/tmp/cosmos3_transfer_depth_trtllm{suffix}").write_bytes(response.content) +``` + +Only edge and blur can be generated from the uploaded source (`"edge": true` +or `"blur": true`); depth, segmentation, and WSM require precomputed control +media. Use `use_guardrails`, not vLLM-Omni's `guardrails`, and send encoded +control bytes rather than a server-local `control_path`. When neither output +dimension is specified, TensorRT-LLM chooses the nearest supported bucket from +the source/control aspect ratio. The checked-in examples explicitly request +1280×720 and their native frame count/fps; WSM uses 101 frames at 10 fps. + +### TensorRT-LLM notebook walkthrough + +[`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) +runs edge, blur, depth, segmentation, and WSM against an already-running Nano +or Super server. It reuses the checked-in media and captions, shows the complete +mode matrix, validates the multipart/base64 request contract, and previews the +synchronous encoded-video responses. + ## Run with vLLM-Omni ### Quickstart @@ -315,6 +392,9 @@ Key fields: - [`run_video_transfer_with_diffusers.ipynb`](./run_video_transfer_with_diffusers.ipynb) — full tutorial for the Diffusers modular pipeline: five single-control transfers on Nano, then the same five on Super, driven by the same specs. +- [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) — + five single-control transfers through TensorRT-LLM's synchronous VisualGen API, + using multipart source uploads and base64 precomputed controls where required. - [`run_video_transfer_with_vllm_omni.ipynb`](./run_video_transfer_with_vllm_omni.ipynb) — full tutorial against an already-running vLLM-Omni server: endpoint checks, repo-local control paths, five single-control transfer requests, and compact previews. The API diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb new file mode 100644 index 00000000..bcab2b11 --- /dev/null +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -0,0 +1,425 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "transfer-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Video Transfer with TensorRT-LLM\n", + "\n", + "Run the five established controls—Canny/edge, blur, depth, segmentation, and WSM—against an\n", + "already-running TensorRT-LLM VisualGen server. The same notebook works with Cosmos3-Nano or\n", + "Cosmos3-Super; select the checkpoint when launching the server.\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then launch one\n", + "server from the TensorRT-LLM checkout:\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "\n", + "# Nano: one GPU\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "\n", + "# Super: four GPUs (run instead of Nano)\n", + "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \\\n", + " nvidia/Cosmos3-Super \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install `requests` in this notebook kernel if needed. The response is synchronous encoded video\n", + "bytes. Inputs use multipart `input_reference`; precomputed depth/seg/WSM controls are base64 media\n", + "inside JSON `extra_params`, because JSON has no byte type. TensorRT-LLM decodes them back to bytes\n", + "at the HTTP boundary.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "TRANSFER_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"transfer\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_TRANSFER_OUTPUT_ROOT\", TRANSFER_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"TRANSFER_ROOT:\", TRANSFER_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-matrix", + "metadata": {}, + "source": [ + "## Request matrix\n", + "\n", + "| Control | Transport | Frames / fps | Text / control guidance |\n", + "| --- | --- | ---: | ---: |\n", + "| Edge | Uploaded source; GPU Canny | 121 / 30 | 3.0 / 1.5 |\n", + "| Blur | Uploaded source; GPU bilateral blur | 121 / 30 | 3.0 / 1.5 |\n", + "| Depth | Uploaded source + base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Segmentation | Uploaded source + base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", + "| WSM | Uploaded source + base64 precomputed control | 101 / 10 | 1.0 / 3.0 |\n", + "\n", + "All checked-in controls are 16:9 and requests explicitly select `1280x720`. TensorRT-LLM can\n", + "also derive the nearest supported bucket from an uploaded source when neither dimension is set.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import base64\n", + "import json\n", + "import time\n", + "\n", + "import requests\n", + "from IPython.display import Video, display\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/generations\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "CONTROL_CASES = {\n", + " \"edge\": {\n", + " \"control\": \"assets/edge/control_edge.mp4\",\n", + " \"prompt\": \"assets/edge/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " \"hint\": {\"preset_edge_threshold\": \"medium\"},\n", + " },\n", + " \"blur\": {\n", + " \"control\": \"assets/blur/control_blur.mp4\",\n", + " \"prompt\": \"assets/blur/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " \"hint\": {\"preset_blur_strength\": \"medium\"},\n", + " },\n", + " \"depth\": {\n", + " \"control\": \"assets/depth/control_depth.mp4\",\n", + " \"prompt\": \"assets/depth/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " },\n", + " \"seg\": {\n", + " \"control\": \"assets/seg/control_seg.mp4\",\n", + " \"prompt\": \"assets/seg/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 2.0,\n", + " },\n", + " \"wsm\": {\n", + " \"control\": \"assets/wsm/control_wsm.mp4\",\n", + " \"prompt\": \"assets/wsm/prompt.json\",\n", + " \"num_frames\": 101,\n", + " \"fps\": 10,\n", + " \"guidance_scale\": 1.0,\n", + " \"control_guidance\": 3.0,\n", + " },\n", + "}\n", + "\n", + "\n", + "def compact_json(path: Path) -> str:\n", + " return json.dumps(json.loads(path.read_text()), ensure_ascii=True, separators=(\",\", \":\"))\n", + "\n", + "\n", + "def run_transfer(control_name: str) -> Path:\n", + " \"\"\"Upload the source as multipart media and carry precomputed controls as base64.\"\"\"\n", + " spec = CONTROL_CASES[control_name]\n", + " control_path = (TRANSFER_ROOT / spec[\"control\"]).resolve()\n", + " prompt_path = (TRANSFER_ROOT / spec[\"prompt\"]).resolve()\n", + " negative_path = TRANSFER_ROOT / \"assets\" / \"negative_prompt.json\"\n", + " for path in (control_path, prompt_path, negative_path):\n", + " if not path.exists():\n", + " raise FileNotFoundError(path)\n", + "\n", + " hint = dict(spec.get(\"hint\", {}))\n", + " # Edge and blur have on-the-fly GPU preprocessors, so the uploaded source is\n", + " # sufficient. Depth, segmentation and WSM require encoded precomputed media.\n", + " if control_name not in {\"edge\", \"blur\"}:\n", + " hint[\"control\"] = base64.b64encode(control_path.read_bytes()).decode(\"ascii\")\n", + "\n", + " extra_params = {\n", + " \"use_duration_template\": False,\n", + " \"use_resolution_template\": False,\n", + " \"use_system_prompt\": False,\n", + " \"use_guardrails\": True,\n", + " control_name: hint,\n", + " \"control_guidance\": spec[\"control_guidance\"],\n", + " \"num_video_frames_per_chunk\": spec[\"num_frames\"],\n", + " \"num_conditional_frames\": 1,\n", + " \"num_first_chunk_conditional_frames\": 0,\n", + " \"max_frames\": spec[\"num_frames\"],\n", + " }\n", + " form = {\n", + " \"prompt\": compact_json(prompt_path),\n", + " \"negative_prompt\": compact_json(negative_path),\n", + " \"size\": \"1280x720\",\n", + " \"num_frames\": str(spec[\"num_frames\"]),\n", + " \"fps\": str(spec[\"fps\"]),\n", + " \"num_inference_steps\": \"35\",\n", + " \"guidance_scale\": str(spec[\"guidance_scale\"]),\n", + " \"max_sequence_length\": \"4096\",\n", + " \"seed\": \"2026\",\n", + " \"format\": \"auto\",\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"video/mp4, video/x-msvideo\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with control_path.open(\"rb\") as source_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"input_reference\": (control_path.name, source_file, \"video/mp4\")},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\").lower()\n", + " suffix = \".avi\" if \"avi\" in content_type or response.content[:4] == b\"RIFF\" else \".mp4\"\n", + " if not response.content:\n", + " raise RuntimeError(\"TensorRT-LLM returned an empty media response\")\n", + " output_path = OUTPUT_ROOT / f\"transfer_{control_name}{suffix}\"\n", + " output_path.write_bytes(response.content)\n", + " print(\"control:\", control_path)\n", + " print(\"extra_params:\", json.dumps({**extra_params, control_name: \"\"}, indent=2))\n", + " print(\"saved:\", output_path, content_type)\n", + " return output_path\n", + "\n", + "\n", + "def view_transfer(path: Path) -> None:\n", + " display(Video(str(path), embed=True))\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-edge-title", + "metadata": {}, + "source": [ + "## Canny / Edge transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-edge-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "edge_output = run_transfer(\"edge\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-edge-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(edge_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-blur-title", + "metadata": {}, + "source": [ + "## Blur transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-blur-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "blur_output = run_transfer(\"blur\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-blur-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(blur_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-depth-title", + "metadata": {}, + "source": [ + "## Depth transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-depth-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "depth_output = run_transfer(\"depth\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-depth-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(depth_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-seg-title", + "metadata": {}, + "source": [ + "## Segmentation transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-seg-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "seg_output = run_transfer(\"seg\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-seg-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(seg_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-wsm-title", + "metadata": {}, + "source": [ + "## World Scenario Model (WSM) transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-wsm-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "wsm_output = run_transfer(\"wsm\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-wsm-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(wsm_output)\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} From a454afddd68bfa4a5e254fb63b8be5b8e2fb880e Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 28 Aug 2026 12:01:01 -0700 Subject: [PATCH 03/14] Update TensorRT-LLM Action and Transfer recipes Signed-off-by: Igor Shovkun --- README.md | 2 +- cookbooks/cosmos3/README.md | 24 ++++--- cookbooks/cosmos3/generator/action/README.md | 9 ++- .../action/run_fd_with_trt_llm.ipynb | 8 +-- .../action/run_id_with_trt_llm.ipynb | 8 +-- .../action/run_policy_with_trt_llm.ipynb | 8 +-- .../cosmos3/generator/transfer/README.md | 66 ++++++++++--------- .../run_video_transfer_with_trt_llm.ipynb | 62 ++++++++--------- 8 files changed, 101 insertions(+), 86 deletions(-) diff --git a/README.md b/README.md index 4ee5568e..bf1d6758 100644 --- a/README.md +++ b/README.md @@ -1155,7 +1155,7 @@ We are building examples that show Cosmos 3 Super/Nano/Edge capabilities end to | Inverse dynamics with SGLang | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | | Action policy with TensorRT-LLM | Generator | DROID policy inference from a concatenated three-camera first frame, returning a rollout and predicted action chunk through the VisualGen API. | [Notebook](cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | | Transfer with Cosmos Framework | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | -| Transfer with TensorRT-LLM | Generator | Edge, blur, depth, segmentation, and WSM transfer using multipart sources, inline precomputed controls, and synchronous encoded-video responses. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | +| Transfer with TensorRT-LLM | Generator | Edge, blur, depth, segmentation, and WSM transfer using inline precomputed controls and synchronous encoded-video responses; raw uploads are used only for server-derived edge/blur. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | | Transfer with vLLM-Omni | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | | Reasoner with Cosmos Framework | Reasoner | Text and image reasoning: detailed captioning, robot task planning, 2D grounding, describe-anything, and action-trajectory prompts, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/reasoner/run_with_cosmos_framework.ipynb) | | Reasoner with vLLM | Reasoner | Image and video reasoning: captioning, temporal localization, embodied reasoning, common-sense reasoning, 2D grounding, describe-anything, action CoT, driving scenes, physical-plausibility, and situation understanding, against an OpenAI-compatible vLLM server (Cosmos3-Super on 4 GPUs by default; switch to Nano per the cookbook README). | [Notebook](cookbooks/cosmos3/reasoner/run_with_vllm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/reasoner/run_with_vllm.ipynb) | diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 0bb705dd..095177d1 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -174,7 +174,9 @@ Action examples, including the published four-step T2I and I2V students. Initial Cosmos3 support was added in TensorRT-LLM PR [#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824), synchronized audio in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and -video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155). +video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155), +Transfer in [#16394](https://github.com/NVIDIA/TensorRT-LLM/pull/16394), and +Action in [#17325](https://github.com/NVIDIA/TensorRT-LLM/pull/17325). Use a TensorRT-LLM checkout or package that includes those changes. Install TensorRT-LLM following its upstream documentation. @@ -264,8 +266,10 @@ trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ Action uses the Nano launch above. For DROID policy inference, replace the model with `nvidia/Cosmos3-Nano-Policy-DROID` and keep the one-GPU Nano config. -The server exposes `/health`, `/v1/videos/generations`, `/v1/videos`, and -`/v1/images/generations`. The audiovisual notebook uses the validated video +The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous +`/v1/videos`, and `/v1/images/generations`. The older +`/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. +The audiovisual notebook uses the validated video generation endpoint for text-to-image, text-to-video, image-to-video, video-to-video, and synchronized audio. Cosmos3 text-to-image is sent as a one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; the notebook @@ -280,12 +284,14 @@ Cosmos3 controls through `extra_params`, so use a TensorRT-LLM build that includ the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. -Transfer uses the synchronous `/v1/videos/generations` route. Upload the source -video as multipart `input_reference`; send `edge`, `blur`, `depth`, `seg`, or -`wsm` under `extra_params`. Precomputed controls are base64-encoded media in the -JSON request and are decoded to bytes at the HTTP boundary. TensorRT-LLM uses -`use_guardrails` for its per-request safety switch; `guardrails`, -`control_path`, and other vLLM-Omni-only names are not interchangeable. +Transfer uses the synchronous `/v1/videos/sync` route. For server-derived edge +or blur, upload the raw source video as multipart `input_reference` and set the +corresponding `extra_params` hint to `true`. For a precomputed edge, blur, +depth, segmentation, or WSM control, base64-encode the control inside its hint; +no `input_reference` is needed. The server decodes inline media to bytes at the +HTTP boundary. TensorRT-LLM uses `use_guardrails` for its per-request safety +switch; `guardrails`, `control_path`, and other vLLM-Omni-only names are not +interchangeable. Action requests use the same synchronous route and upload an image or video as `input_reference`. Because an action trajectory cannot be represented in MP4 or diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index b3abecc7..31189097 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -191,10 +191,11 @@ actions = json.loads((action_root / "assets/actions/av_traj_forward.json").read_ with image_path.open("rb") as image_file: response = requests.post( - "http://localhost:8000/v1/videos/generations", + "http://localhost:8000/v1/videos/sync", data={ "prompt": "You are an autonomous vehicle planning system.", "format": "safetensors", + "response_format": "file", "seed": "0", "extra_params": json.dumps( { @@ -222,8 +223,10 @@ Action modes therefore use a tensor output: `format=auto` selects `safetensors`, and explicit `mp4`/`avi` is rejected rather than dropping the trajectory. -The quickstart uses the blocking `POST /v1/videos/generations` route. The -asynchronous `POST /v1/videos` route returns HTTP 202; poll +The quickstart uses the blocking `POST /v1/videos/sync` route with +`response_format=file`, which returns the tensor payload bytes directly. The +older `/v1/videos/generations` spelling is a deprecated alias. The asynchronous +`POST /v1/videos` route returns HTTP 202; poll `GET /v1/videos/{id}` and download the same tensor payload from `GET /v1/videos/{id}/content`. diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb index 62158348..6c8d7dbc 100644 --- a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -85,8 +85,8 @@ "source": [ "## TensorRT-LLM response contract\n", "\n", - "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", - "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", + "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", "\n", "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", @@ -117,7 +117,7 @@ "\n", "def video_api_url() -> str:\n", " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", - " return f\"{root}/videos/generations\"\n", + " return f\"{root}/videos/sync\"\n", "\n", "\n", "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", @@ -145,7 +145,7 @@ " form = {\n", " \"prompt\": prompt,\n", " \"format\": \"safetensors\",\n", - " \"response_format\": \"url\",\n", + " \"response_format\": \"file\",\n", " \"seed\": \"0\",\n", " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", " }\n", diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb index 6258ee4d..2bf3cd29 100644 --- a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -84,8 +84,8 @@ "source": [ "## TensorRT-LLM response contract\n", "\n", - "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", - "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", + "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", "\n", "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", @@ -116,7 +116,7 @@ "\n", "def video_api_url() -> str:\n", " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", - " return f\"{root}/videos/generations\"\n", + " return f\"{root}/videos/sync\"\n", "\n", "\n", "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", @@ -144,7 +144,7 @@ " form = {\n", " \"prompt\": prompt,\n", " \"format\": \"safetensors\",\n", - " \"response_format\": \"url\",\n", + " \"response_format\": \"file\",\n", " \"seed\": \"0\",\n", " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", " }\n", diff --git a/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb index f0021dc0..22143c23 100644 --- a/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb @@ -84,8 +84,8 @@ "source": [ "## TensorRT-LLM response contract\n", "\n", - "The notebooks use the blocking `POST /v1/videos/generations` route. Every Action request sets\n", - "`format=safetensors`; the response is `application/octet-stream` with named `video`, `action`,\n", + "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", + "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", "\n", "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", @@ -116,7 +116,7 @@ "\n", "def video_api_url() -> str:\n", " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", - " return f\"{root}/videos/generations\"\n", + " return f\"{root}/videos/sync\"\n", "\n", "\n", "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", @@ -144,7 +144,7 @@ " form = {\n", " \"prompt\": prompt,\n", " \"format\": \"safetensors\",\n", - " \"response_format\": \"url\",\n", + " \"response_format\": \"file\",\n", " \"seed\": \"0\",\n", " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", " }\n", diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index f5778a9e..ea8fbc25 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -25,8 +25,9 @@ Environment setup is centralized in the shared Video transfer generates a target clip from a `prompt.json` caption and one or more spatial control signals. The Framework path uses `model_mode` `video2video` in a local JSON spec. -TensorRT-LLM uses `POST /v1/videos/generations`: source media is multipart -`input_reference`, while hints are carried under `extra_params`. +TensorRT-LLM uses `POST /v1/videos/sync`. A raw source used to derive edge or +blur is multipart `input_reference`; precomputed controls are carried directly +under their `extra_params` hint and do not need an `input_reference`. The vLLM-Omni path uses `POST /v1/videos/sync` and passes one or more hint keys (`edge`, `blur`, `depth`, `seg`, or `wsm`) inside `extra_params`. Cosmos Framework accepts pre-computed control videos (`control_path`) or derives active controls from a raw source video @@ -199,9 +200,11 @@ reuses the previews from [`preview_helpers.py`](./preview_helpers.py), writing o Set up and launch Nano or Super as described in the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). TensorRT-LLM -accepts source media as multipart `input_reference`. JSON cannot carry bytes, -so a precomputed control is base64 encoded inside its hint object; the server -decodes it at the HTTP boundary before validating and dispatching the request. +accepts a raw source video as multipart `input_reference` when edge or blur will +be computed on the server. The checked-in assets are already precomputed +controls, so this example base64-encodes the control inside its hint and does +not send an `input_reference`. The server decodes inline media at the HTTP +boundary before validating and dispatching the request. ```python import base64 @@ -229,36 +232,39 @@ extra_params = { "max_frames": 121, } -with control_path.open("rb") as source_file: - response = requests.post( - "http://localhost:8000/v1/videos/generations", - data={ - "prompt": prompt, - "negative_prompt": negative, - "size": "1280x720", - "num_frames": "121", - "fps": "30", - "num_inference_steps": "35", - "guidance_scale": "3.0", - "max_sequence_length": "4096", - "seed": "2026", - "extra_params": json.dumps(extra_params), - }, - files={"input_reference": (control_path.name, source_file, "video/mp4")}, - headers={"Accept": "video/mp4, video/x-msvideo"}, - ) +response = requests.post( + "http://localhost:8000/v1/videos/sync", + json={ + "prompt": prompt, + "negative_prompt": negative, + "size": "1280x720", + "num_frames": 121, + "fps": 30, + "num_inference_steps": 35, + "guidance_scale": 3.0, + "max_sequence_length": 4096, + "seed": 2026, + "format": "auto", + "response_format": "file", + "extra_params": extra_params, + }, + headers={"Accept": "video/mp4, video/x-msvideo"}, +) response.raise_for_status() suffix = ".avi" if "avi" in response.headers.get("content-type", "") else ".mp4" Path(f"/tmp/cosmos3_transfer_depth_trtllm{suffix}").write_bytes(response.content) ``` -Only edge and blur can be generated from the uploaded source (`"edge": true` -or `"blur": true`); depth, segmentation, and WSM require precomputed control -media. Use `use_guardrails`, not vLLM-Omni's `guardrails`, and send encoded -control bytes rather than a server-local `control_path`. When neither output -dimension is specified, TensorRT-LLM chooses the nearest supported bucket from -the source/control aspect ratio. The checked-in examples explicitly request -1280×720 and their native frame count/fps; WSM uses 101 frames at 10 fps. +Only edge and blur can be generated from a raw uploaded source (`"edge": true` +or `"blur": true`); depth, segmentation, and WSM always require precomputed +control media. Precomputed edge and blur are also accepted. Use +`use_guardrails`, not vLLM-Omni's `guardrails`, and send encoded control bytes +rather than a server-local `control_path`. When neither output dimension is +specified, TensorRT-LLM chooses the nearest supported bucket from the source or +first precomputed control's aspect ratio. The checked-in examples explicitly +request 1280×720 and their native frame count/fps; WSM uses 101 frames at 10 +fps. `/v1/videos/generations` remains only as a deprecated alias of the +canonical blocking `/v1/videos/sync` route. ### TensorRT-LLM notebook walkthrough diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb index bcab2b11..8d042a98 100644 --- a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -47,9 +47,9 @@ "```\n", "\n", "Install `requests` in this notebook kernel if needed. The response is synchronous encoded video\n", - "bytes. Inputs use multipart `input_reference`; precomputed depth/seg/WSM controls are base64 media\n", - "inside JSON `extra_params`, because JSON has no byte type. TensorRT-LLM decodes them back to bytes\n", - "at the HTTP boundary.\n" + "bytes. Every checked-in asset is already a precomputed control, so the notebook sends it once as\n", + "base64 media inside JSON `extra_params`; no `input_reference` upload is needed. TensorRT-LLM\n", + "decodes the media back to bytes at the HTTP boundary.\n" ] }, { @@ -94,14 +94,16 @@ "\n", "| Control | Transport | Frames / fps | Text / control guidance |\n", "| --- | --- | ---: | ---: |\n", - "| Edge | Uploaded source; GPU Canny | 121 / 30 | 3.0 / 1.5 |\n", - "| Blur | Uploaded source; GPU bilateral blur | 121 / 30 | 3.0 / 1.5 |\n", - "| Depth | Uploaded source + base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", - "| Segmentation | Uploaded source + base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", - "| WSM | Uploaded source + base64 precomputed control | 101 / 10 | 1.0 / 3.0 |\n", + "| Edge | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Blur | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Depth | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Segmentation | Base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", + "| WSM | Base64 precomputed control | 101 / 10 | 1.0 / 3.0 |\n", "\n", "All checked-in controls are 16:9 and requests explicitly select `1280x720`. TensorRT-LLM can\n", - "also derive the nearest supported bucket from an uploaded source when neither dimension is set.\n" + "also derive the nearest supported bucket from the first precomputed control when neither dimension\n", + "is set. To compute edge or blur on the server instead, upload a raw source as multipart\n", + "`input_reference` and set the corresponding hint to `true`.\n" ] }, { @@ -125,7 +127,7 @@ "\n", "def video_api_url() -> str:\n", " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", - " return f\"{root}/videos/generations\"\n", + " return f\"{root}/videos/sync\"\n", "\n", "\n", "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", @@ -194,7 +196,7 @@ "\n", "\n", "def run_transfer(control_name: str) -> Path:\n", - " \"\"\"Upload the source as multipart media and carry precomputed controls as base64.\"\"\"\n", + " \"\"\"Send one precomputed control through the JSON/base64 input contract.\"\"\"\n", " spec = CONTROL_CASES[control_name]\n", " control_path = (TRANSFER_ROOT / spec[\"control\"]).resolve()\n", " prompt_path = (TRANSFER_ROOT / spec[\"prompt\"]).resolve()\n", @@ -204,10 +206,9 @@ " raise FileNotFoundError(path)\n", "\n", " hint = dict(spec.get(\"hint\", {}))\n", - " # Edge and blur have on-the-fly GPU preprocessors, so the uploaded source is\n", - " # sufficient. Depth, segmentation and WSM require encoded precomputed media.\n", - " if control_name not in {\"edge\", \"blur\"}:\n", - " hint[\"control\"] = base64.b64encode(control_path.read_bytes()).decode(\"ascii\")\n", + " # These files are already edge/blur/depth/seg/WSM control representations.\n", + " # Send each one directly instead of also treating it as a raw input video.\n", + " hint[\"control\"] = base64.b64encode(control_path.read_bytes()).decode(\"ascii\")\n", "\n", " extra_params = {\n", " \"use_duration_template\": False,\n", @@ -221,30 +222,29 @@ " \"num_first_chunk_conditional_frames\": 0,\n", " \"max_frames\": spec[\"num_frames\"],\n", " }\n", - " form = {\n", + " request_body = {\n", " \"prompt\": compact_json(prompt_path),\n", " \"negative_prompt\": compact_json(negative_path),\n", " \"size\": \"1280x720\",\n", - " \"num_frames\": str(spec[\"num_frames\"]),\n", - " \"fps\": str(spec[\"fps\"]),\n", - " \"num_inference_steps\": \"35\",\n", - " \"guidance_scale\": str(spec[\"guidance_scale\"]),\n", - " \"max_sequence_length\": \"4096\",\n", - " \"seed\": \"2026\",\n", + " \"num_frames\": spec[\"num_frames\"],\n", + " \"fps\": spec[\"fps\"],\n", + " \"num_inference_steps\": 35,\n", + " \"guidance_scale\": spec[\"guidance_scale\"],\n", + " \"max_sequence_length\": 4096,\n", + " \"seed\": 2026,\n", " \"format\": \"auto\",\n", - " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " \"response_format\": \"file\",\n", + " \"extra_params\": extra_params,\n", " }\n", " headers = {\"Accept\": \"video/mp4, video/x-msvideo\"}\n", " if TRTLLM_API_KEY:\n", " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", - " with control_path.open(\"rb\") as source_file:\n", - " response = requests.post(\n", - " video_api_url(),\n", - " data=form,\n", - " files={\"input_reference\": (control_path.name, source_file, \"video/mp4\")},\n", - " headers=headers,\n", - " timeout=3600,\n", - " )\n", + " response = requests.post(\n", + " video_api_url(),\n", + " json=request_body,\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", " if not response.ok:\n", " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", " content_type = response.headers.get(\"content-type\", \"\").lower()\n", From 9ceef65ba65a4af0efdba2a3a277fac45e753d9d Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 28 Aug 2026 23:08:21 -0700 Subject: [PATCH 04/14] Fix TensorRT-LLM guardrail cookbook setup Signed-off-by: Igor Shovkun --- cookbooks/cosmos3/README.md | 6 ++++-- .../cosmos3/generator/transfer/assets/depth/prompt.json | 4 ++-- .../generator/transfer/assets/negative_prompt.json | 8 ++++---- 3 files changed, 10 insertions(+), 8 deletions(-) diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 095177d1..27a25f90 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -220,8 +220,10 @@ explicitly disable guardrails before starting the server: ```bash pip install cosmos_guardrail==0.3.0 -# If needed by your OpenCV stack: -# pip uninstall opencv-python +# On headless servers without libGL.so.1, replace the OpenCV wheel pulled in by +# cosmos_guardrail with the matching headless build: +pip uninstall -y opencv-python +pip install opencv-python-headless==5.0.0.93 ``` Set the TensorRT-LLM source root for the shared VisualGen config YAMLs: diff --git a/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json b/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json index 90f2dd01..59ba72f3 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json @@ -62,7 +62,7 @@ }, { "description": "A white sedan driving slowly forward in the center of the road", - "appearance_details": "Clean white paint, standard compact sedan body, brake lights occasionally glowing red", + "appearance_details": "Clean white paint, standard compact sedan body, rear lights occasionally glowing red", "relationship": "Directly ahead of the observer, pacing the forward motion", "location": "Center foreground to midground", "relative_size": "Medium within frame", @@ -198,4 +198,4 @@ "aspect_ratio": "16,9", "duration": "4s", "fps": 30 -} \ No newline at end of file +} diff --git a/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json b/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json index 44dff269..0b34964d 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json @@ -2,7 +2,7 @@ "subjects": [ { "description": "Blurry, poorly defined subjects with inconsistent shapes and unrealistic proportions.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -22,7 +22,7 @@ }, { "description": "Extremely low-quality subjects with visible rendering artifacts, broken mesh geometry, and completely unrealistic proportions throughout.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -42,7 +42,7 @@ }, { "description": "Poorly generated subjects exhibiting all hallmarks of failed neural rendering \u2014 flickering edges, inconsistent depth, and uncanny spatial relationships.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -71,7 +71,7 @@ "aesthetics": { "composition": "Cluttered, poorly framed composition with no clear focal point. Important elements are cut off by the frame edges. The rule of thirds is completely ignored, leading to an unbalanced and visually unpleasant arrangement.", "color_scheme": "Oversaturated, garish colors that clash violently. Color banding is visible in gradient areas. The overall palette feels artificial and digitally processed rather than natural.", - "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels lifeless and sterile despite attempting to portray dynamic action.", + "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels flat and sterile despite attempting to portray dynamic action.", "patterns": "Visible tiling artifacts in textures, moir\u00e9 patterns, and aliasing on edges." }, "cinematography": { From bf2c652f6edd7ca9c73df89d894c76ffe436cece Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Mon, 31 Aug 2026 15:21:33 -0700 Subject: [PATCH 05/14] docs: align TRT-LLM cookbooks with runtime validation Signed-off-by: Igor Shovkun --- README.md | 1 - cookbooks/cosmos3/README.md | 3 - cookbooks/cosmos3/generator/action/README.md | 11 +- .../action/run_policy_with_trt_llm.ipynb | 324 ------------------ .../image2video/neg_prompt.json | 8 +- .../prompts/image2video/humanoid_robot.json | 4 +- .../cosmos3/generator/transfer/README.md | 4 +- .../transfer/assets/edge/prompt.json | 4 +- .../run_video_transfer_with_trt_llm.ipynb | 4 +- .../cosmos3/generator/transfer/specs/wsm.json | 4 +- 10 files changed, 18 insertions(+), 349 deletions(-) delete mode 100644 cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb diff --git a/README.md b/README.md index ca4aa569..555e48fe 100644 --- a/README.md +++ b/README.md @@ -1170,7 +1170,6 @@ We are building examples that show Cosmos 3 Super/Nano/Edge capabilities end to | Inverse dynamics with TensorRT-LLM | Generator | Inverse dynamics from a checked-in AV observation clip, returning reconstructed video and the predicted trajectory in a tensor payload. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb) | | Inverse dynamics with vLLM-Omni | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_vllm_omni.ipynb) | | Inverse dynamics with SGLang | Generator | Inverse dynamics: ego-motion trajectory prediction from input AV video, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_id_with_sglang.ipynb) | -| Action policy with TensorRT-LLM | Generator | DROID policy inference from a concatenated three-camera first frame, returning a rollout and predicted action chunk through the VisualGen API. | [Notebook](cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb) | | Transfer with Cosmos Framework | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_cosmos_framework.ipynb) | | Transfer with TensorRT-LLM | Generator | Edge, blur, depth, segmentation, and WSM transfer using inline precomputed controls and synchronous encoded-video responses; raw uploads are used only for server-derived edge/blur. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb) | | Transfer with vLLM-Omni | Generator | Video transfer: edge, blur, depth, segmentation, and world-scenario controls with captions, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_vllm_omni.ipynb) | diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index e6e68478..137d7945 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -265,9 +265,6 @@ trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ --port "$COSMOS3_TRTLLM_PORT" ``` -Action uses the Nano launch above. For DROID policy inference, replace the model -with `nvidia/Cosmos3-Nano-Policy-DROID` and keep the one-GPU Nano config. - The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous `/v1/videos`, and `/v1/images/generations`. The older `/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 31189097..059a1d0d 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -217,9 +217,9 @@ print(payload["video"].shape, payload["action"].shape, payload["frame_rate"].ite The `av` domain preset supplies its 60-step, 9D, 480p/10 fps recipe; other recognized domains similarly fill omitted action width, chunk, resolution, and -frame rate. Policy and inverse dynamics predict their action tensors, while -forward dynamics carries the conditioned action alongside the rollout. All -Action modes therefore use a tensor output: `format=auto` selects +frame rate. Inverse dynamics predicts its action tensor, while forward dynamics +carries the conditioned action alongside the rollout. Both supported modes use +a tensor output: `format=auto` selects `safetensors`, and explicit `mp4`/`avi` is rejected rather than dropping the trajectory. @@ -236,11 +236,8 @@ older `/v1/videos/generations` spelling is a deprecated alias. The asynchronous dynamics from the checked-in AV image and 60×9 action trajectory. - [`run_id_with_trt_llm.ipynb`](./run_id_with_trt_llm.ipynb) — inverse dynamics from the checked-in AV observation clip. -- [`run_policy_with_trt_llm.ipynb`](./run_policy_with_trt_llm.ipynb) — DROID - policy inference with a concatenated three-camera first frame and - `nvidia/Cosmos3-Nano-Policy-DROID`. -Each notebook decodes TensorRT-LLM's `safetensors` response, writes the rollout +Both notebooks decode TensorRT-LLM's `safetensors` response, write the rollout video, and inspects the named `action` tensor. This response contract is not the top-level action JSON used by vLLM-Omni. diff --git a/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb deleted file mode 100644 index 22143c23..00000000 --- a/cookbooks/cosmos3/generator/action/run_policy_with_trt_llm.ipynb +++ /dev/null @@ -1,324 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "id": "policy-trt-license", - "metadata": {}, - "source": [ - "\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-title", - "metadata": {}, - "source": [ - "# Cosmos3 Action Policy with TensorRT-LLM\n", - "\n", - "Build the trained DROID concatenated camera view from checked-in clips, then jointly predict a\n", - "16-step, 10D action chunk and a 17-frame rollout with `Cosmos3-Nano-Policy-DROID`.\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-server", - "metadata": {}, - "source": [ - "## Start the TensorRT-LLM server\n", - "\n", - "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", - "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", - "\n", - "```bash\n", - "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", - "trtllm-serve nvidia/Cosmos3-Nano-Policy-DROID \\\n", - " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", - " --port 8000\n", - "```\n", - "\n", - "Install client-side decoding dependencies in this notebook kernel if needed:\n", - "\n", - "```bash\n", - "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", - "```\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "policy-trt-setup", - "metadata": {}, - "outputs": [], - "source": [ - "from pathlib import Path\n", - "import os\n", - "\n", - "\n", - "def find_repo_root(start: Path) -> Path:\n", - " for path in [start, *start.parents]:\n", - " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", - " return path\n", - " return start\n", - "\n", - "\n", - "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", - "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", - "OUTPUT_ROOT = Path(\n", - " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", - ").resolve()\n", - "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", - "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", - "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", - "\n", - "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", - "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", - "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", - "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-contract", - "metadata": {}, - "source": [ - "## TensorRT-LLM response contract\n", - "\n", - "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", - "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", - "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", - "\n", - "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", - "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", - "`GET /v1/videos/{id}/content`.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "policy-trt-helpers", - "metadata": {}, - "outputs": [], - "source": [ - "import json\n", - "import mimetypes\n", - "import time\n", - "\n", - "import imageio.v3 as iio\n", - "import requests\n", - "from IPython.display import Video, display\n", - "from safetensors.torch import load as load_safetensors\n", - "\n", - "\n", - "def server_root_url() -> str:\n", - " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", - "\n", - "\n", - "def video_api_url() -> str:\n", - " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", - " return f\"{root}/videos/sync\"\n", - "\n", - "\n", - "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", - " deadline = time.time() + timeout_s\n", - " url = f\"{server_root_url()}/health\"\n", - " while time.time() < deadline:\n", - " try:\n", - " response = requests.get(url, timeout=10)\n", - " if response.ok:\n", - " print(\"TensorRT-LLM server is ready:\", url)\n", - " return\n", - " except requests.RequestException as exc:\n", - " print(\"waiting for TensorRT-LLM:\", exc)\n", - " time.sleep(interval_s)\n", - " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", - "\n", - "\n", - "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", - " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", - " reference_path = reference_path.resolve()\n", - " if not reference_path.exists():\n", - " raise FileNotFoundError(reference_path)\n", - " output_path.parent.mkdir(parents=True, exist_ok=True)\n", - " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", - " form = {\n", - " \"prompt\": prompt,\n", - " \"format\": \"safetensors\",\n", - " \"response_format\": \"file\",\n", - " \"seed\": \"0\",\n", - " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", - " }\n", - " headers = {\"Accept\": \"application/octet-stream\"}\n", - " if TRTLLM_API_KEY:\n", - " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", - " with reference_path.open(\"rb\") as reference_file:\n", - " response = requests.post(\n", - " video_api_url(),\n", - " data=form,\n", - " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", - " headers=headers,\n", - " timeout=3600,\n", - " )\n", - " if not response.ok:\n", - " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", - " content_type = response.headers.get(\"content-type\", \"\")\n", - " if \"application/octet-stream\" not in content_type:\n", - " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", - " output_path.write_bytes(response.content)\n", - " payload = load_safetensors(response.content)\n", - " missing = {\"video\", \"action\"} - payload.keys()\n", - " if missing:\n", - " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", - " print(\"saved tensor payload:\", output_path)\n", - " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", - " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", - " return payload\n", - "\n", - "\n", - "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", - " frames = payload[\"video\"].detach().cpu().numpy()\n", - " if frames.ndim == 5 and frames.shape[0] == 1:\n", - " frames = frames[0]\n", - " if frames.ndim != 4 or frames.shape[-1] != 3:\n", - " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", - " fps_tensor = payload.get(\"frame_rate\")\n", - " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", - " iio.imwrite(output_path, frames, fps=fps)\n", - " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", - " display(Video(str(output_path), embed=True))\n", - " return output_path\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-input-title", - "metadata": {}, - "source": [ - "## Build the DROID multiview first frame\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "policy-trt-input", - "metadata": {}, - "outputs": [], - "source": [ - "import subprocess\n", - "\n", - "import imageio_ffmpeg\n", - "from IPython.display import Image as NotebookImage\n", - "from PIL import Image, ImageOps\n", - "\n", - "DROID_ROOT = ACTION_ROOT / \"assets\" / \"droid_lerobot_example\"\n", - "camera_paths = {\n", - " \"wrist\": DROID_ROOT / \"videos/observation.image.wrist_image_left/chunk-000/file-000.mp4\",\n", - " \"left\": DROID_ROOT / \"videos/observation.image.exterior_image_1_left/chunk-000/file-000.mp4\",\n", - " \"right\": DROID_ROOT / \"videos/observation.image.exterior_image_2_left/chunk-000/file-000.mp4\",\n", - "}\n", - "for path in camera_paths.values():\n", - " if not path.exists():\n", - " raise FileNotFoundError(path)\n", - "\n", - "frame_dir = OUTPUT_ROOT / \"policy_inputs\"\n", - "frame_dir.mkdir(parents=True, exist_ok=True)\n", - "ffmpeg = imageio_ffmpeg.get_ffmpeg_exe()\n", - "frames = {}\n", - "for name, video_path in camera_paths.items():\n", - " frame_path = frame_dir / f\"{name}.png\"\n", - " subprocess.run(\n", - " [ffmpeg, \"-y\", \"-loglevel\", \"error\", \"-i\", str(video_path), \"-frames:v\", \"1\", str(frame_path)],\n", - " check=True,\n", - " )\n", - " frames[name] = Image.open(frame_path).convert(\"RGB\")\n", - "\n", - "target_w, target_h = 640, 540\n", - "top_h = target_h // 2\n", - "bottom_h = target_h - top_h\n", - "half_w = target_w // 2\n", - "wrist = ImageOps.fit(frames[\"wrist\"], (target_w, top_h), method=Image.Resampling.BICUBIC)\n", - "left = ImageOps.fit(frames[\"left\"], (half_w, bottom_h), method=Image.Resampling.BICUBIC)\n", - "right = ImageOps.fit(frames[\"right\"], (half_w, bottom_h), method=Image.Resampling.BICUBIC)\n", - "conditioning = Image.new(\"RGB\", (target_w, target_h))\n", - "conditioning.paste(wrist, (0, 0))\n", - "conditioning.paste(left, (0, top_h))\n", - "conditioning.paste(right, (half_w, top_h))\n", - "policy_image_path = frame_dir / \"droid_policy_first_frame.png\"\n", - "conditioning.save(policy_image_path)\n", - "display(NotebookImage(filename=str(policy_image_path), width=640))\n", - "\n", - "prompt = os.environ.get(\n", - " \"COSMOS3_POLICY_PROMPT\",\n", - " \"Pick up the object and place it in the target container.\",\n", - ")\n", - "extra_params = {\n", - " \"action_mode\": \"policy\",\n", - " \"domain_name\": \"droid_lerobot\",\n", - " \"view_point\": \"concat_view\",\n", - " \"use_guardrails\": True,\n", - "}\n", - "print(\"prompt:\", prompt)\n", - "print(json.dumps(extra_params, indent=2))\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-run-title", - "metadata": {}, - "source": [ - "## Run policy inference\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "policy-trt-run", - "metadata": {}, - "outputs": [], - "source": [ - "wait_for_server()\n", - "policy_payload = submit_action(\n", - " prompt,\n", - " policy_image_path,\n", - " extra_params,\n", - " OUTPUT_ROOT / \"policy_droid.safetensors\",\n", - ")\n" - ] - }, - { - "cell_type": "markdown", - "id": "policy-trt-view-title", - "metadata": {}, - "source": [ - "## Inspect the rollout and predicted action chunk\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "policy-trt-view", - "metadata": {}, - "outputs": [], - "source": [ - "save_and_view_video(policy_payload, OUTPUT_ROOT / \"policy_droid.mp4\")\n", - "policy_action = policy_payload[\"action\"].detach().cpu().numpy()\n", - "print(\"predicted action shape:\", policy_action.shape)\n", - "print(\"first five predicted action rows:\")\n", - "print(policy_action[:5])\n" - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": "Python 3", - "language": "python", - "name": "python3" - }, - "language_info": { - "name": "python", - "version": "3" - } - }, - "nbformat": 4, - "nbformat_minor": 5 -} diff --git a/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json b/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json index 2a0bd29b..301387bf 100644 --- a/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json +++ b/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json @@ -2,7 +2,7 @@ "subjects": [ { "description": "Blurry, poorly defined subjects with inconsistent shapes and unrealistic proportions.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -22,7 +22,7 @@ }, { "description": "Extremely low-quality subjects with visible rendering artifacts, broken mesh geometry, and completely unrealistic proportions throughout.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -42,7 +42,7 @@ }, { "description": "Poorly generated subjects exhibiting all hallmarks of failed neural rendering \u2014 flickering edges, inconsistent depth, and uncanny spatial relationships.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -71,7 +71,7 @@ "aesthetics": { "composition": "Cluttered, poorly framed composition with no clear focal point. Important elements are cut off by the frame edges. The rule of thirds is completely ignored, leading to an unbalanced and visually unpleasant arrangement.", "color_scheme": "Oversaturated, garish colors that clash violently. Color banding is visible in gradient areas. The overall palette feels artificial and digitally processed rather than natural.", - "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels lifeless and sterile despite attempting to portray dynamic action.", + "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels flat and sterile despite attempting to portray dynamic action.", "patterns": "Visible tiling artifacts in textures, moir\u00e9 patterns, and aliasing on edges." }, "cinematography": { diff --git a/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json b/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json index 6a14ef50..24f2d76b 100644 --- a/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json +++ b/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json @@ -92,7 +92,7 @@ } ], "transitions": [], - "temporal_caption": "The video opens with a humanoid robot standing motionless in a modern living room. At around one second, the robot begins to shift its weight and bend its knees slightly. By two seconds, it drops into a deep crouch with arms swinging back. At approximately three seconds, the robot explosively launches itself upward and backward, its white body arcing through the air as it tucks its legs tightly. Between three and five seconds, the robot completes a full backward rotation high above the living room floor, its mechanical joints visible as it spins. At five seconds, the robot's feet come back down and make contact with the rug, knees deeply bent to absorb the landing impact. From five to seven seconds, the robot smoothly extends back to its full standing height, returning to a relaxed upright pose as if nothing extraordinary just occurred.", + "temporal_caption": "The video opens with a humanoid robot standing motionless in a modern living room. At around one second, the robot begins to shift its weight and bend its knees slightly. By two seconds, it drops into a deep crouch with arms swinging back. At approximately three seconds, the robot explosively launches itself upward and backward, its white body arcing through the air as it tucks its legs tightly. Between three and five seconds, the robot completes a full backward rotation high above the living room floor, its mechanical joints visible as it spins. At five seconds, the robot's feet come back down and make contact with the rug, knees deeply bent to absorb the landing impact. From 5 to 7 seconds, the robot smoothly extends back to its full standing height, returning to a relaxed upright pose as if nothing extraordinary just occurred.", "resolution": { "W": 1280, "H": 720 @@ -100,4 +100,4 @@ "aspect_ratio": "16,9", "duration": "7s", "fps": 24 - } \ No newline at end of file + } diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index ea8fbc25..350c3840 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -43,7 +43,7 @@ for the negative caption. | Blur | `assets/blur/` | `control_blur.mp4` + `prompt.json` | 121 frames @ 30 FPS | | Depth | `assets/depth/` | `control_depth.mp4` + `prompt.json` | 121 frames @ 30 FPS | | Segmentation | `assets/seg/` | `control_seg.mp4` + `prompt.json` | 121 frames @ 30 FPS | -| World scenario (WSM) | `assets/wsm/` | `control_wsm.mp4` + `prompt.json` | 101 frames @ 10 FPS | +| World scenario (WSM) | `assets/wsm/` | `control_wsm.mp4` + `prompt.json` | 100 frames @ 10 FPS | | Multi-control | `assets/multi_control/` | `vision_path` + multiple hints (Framework example) | 121 frames @ 30 FPS | Transfer inference is selected automatically when any hint key is present in the @@ -262,7 +262,7 @@ control media. Precomputed edge and blur are also accepted. Use rather than a server-local `control_path`. When neither output dimension is specified, TensorRT-LLM chooses the nearest supported bucket from the source or first precomputed control's aspect ratio. The checked-in examples explicitly -request 1280×720 and their native frame count/fps; WSM uses 101 frames at 10 +request 1280×720 and their native frame count/fps; WSM uses 100 frames at 10 fps. `/v1/videos/generations` remains only as a deprecated alias of the canonical blocking `/v1/videos/sync` route. diff --git a/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json b/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json index ec91950b..efdaee76 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json @@ -21,7 +21,7 @@ "number_of_legs": 2 } ], - "background_setting": "A brightly lit rehearsal studio with light-colored wood-style flooring marked by circular green decals used as position markers. One wall is covered by a large black curtain decorated with scattered small white rectangular papers, in front of which two tripods hold small recording cameras. The opposite wall features a large window with bright red frames through which natural daylight streams, revealing a brick building across the way. Beneath the window rests a long bench with alternating orange and white patterned cushions, and a black floor-standing speaker stands nearby.", + "background_setting": "A brightly lit rehearsal studio with light-colored wood-style flooring marked by circular green decals used as position markers. One wall is covered by a large black curtain decorated with scattered small white rectangular papers, in front of which two tripods hold small recording cameras. The opposite wall features a large window with bright red frames through which natural daylight streams, revealing a masonry building across the way. Beneath the window rests a long bench with alternating orange and white patterned cushions, and a black floor-standing speaker stands nearby.", "lighting": { "conditions": "Bright daylight combined with even studio lighting", "direction": "Side-lit from the right through the large red-framed window, with soft ambient fill across the room", @@ -83,4 +83,4 @@ "aspect_ratio": "16,9", "duration": "4s", "fps": 30 -} \ No newline at end of file +} diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb index 8d042a98..1b6c4048 100644 --- a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -98,7 +98,7 @@ "| Blur | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", "| Depth | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", "| Segmentation | Base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", - "| WSM | Base64 precomputed control | 101 / 10 | 1.0 / 3.0 |\n", + "| WSM | Base64 precomputed control | 100 / 10 | 1.0 / 3.0 |\n", "\n", "All checked-in controls are 16:9 and requests explicitly select `1280x720`. TensorRT-LLM can\n", "also derive the nearest supported bucket from the first precomputed control when neither dimension\n", @@ -183,7 +183,7 @@ " \"wsm\": {\n", " \"control\": \"assets/wsm/control_wsm.mp4\",\n", " \"prompt\": \"assets/wsm/prompt.json\",\n", - " \"num_frames\": 101,\n", + " \"num_frames\": 100,\n", " \"fps\": 10,\n", " \"guidance_scale\": 1.0,\n", " \"control_guidance\": 3.0,\n", diff --git a/cookbooks/cosmos3/generator/transfer/specs/wsm.json b/cookbooks/cosmos3/generator/transfer/specs/wsm.json index eb18d535..e56a7b60 100644 --- a/cookbooks/cosmos3/generator/transfer/specs/wsm.json +++ b/cookbooks/cosmos3/generator/transfer/specs/wsm.json @@ -3,12 +3,12 @@ "model_mode": "video2video", "resolution": "720", "aspect_ratio": "16,9", - "num_frames": 101, + "num_frames": 100, "fps": 10, "shift": 10.0, "num_steps": 50, "seed": 2026, - "num_video_frames_per_chunk": 101, + "num_video_frames_per_chunk": 100, "num_conditional_frames": 1, "num_first_chunk_conditional_frames": 0, "share_vision_temporal_positions": true, From 82ac4007761cd64434707592b91e6d75dc795faa Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Mon, 31 Aug 2026 16:01:10 -0700 Subject: [PATCH 06/14] docs: align TRT-LLM Transfer input description Signed-off-by: Igor Shovkun --- cookbooks/cosmos3/generator/transfer/README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index 350c3840..f535f400 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -271,7 +271,7 @@ canonical blocking `/v1/videos/sync` route. [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) runs edge, blur, depth, segmentation, and WSM against an already-running Nano or Super server. It reuses the checked-in media and captions, shows the complete -mode matrix, validates the multipart/base64 request contract, and previews the +mode matrix, validates the JSON/base64 control request contract, and previews the synchronous encoded-video responses. ## Run with vLLM-Omni @@ -400,7 +400,7 @@ Key fields: then the same five on Super, driven by the same specs. - [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) — five single-control transfers through TensorRT-LLM's synchronous VisualGen API, - using multipart source uploads and base64 precomputed controls where required. + using base64-encoded precomputed controls in JSON `extra_params`. - [`run_video_transfer_with_vllm_omni.ipynb`](./run_video_transfer_with_vllm_omni.ipynb) — full tutorial against an already-running vLLM-Omni server: endpoint checks, repo-local control paths, five single-control transfer requests, and compact previews. The API From b367242b870487d6a411fb7019cf6bcaac4473c8 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Wed, 2 Sep 2026 10:27:46 -0700 Subject: [PATCH 07/14] docs: use canonical TensorRT-LLM video endpoint --- cookbooks/cosmos3/README.md | 19 ++++++++++--------- cookbooks/cosmos3/generator/action/README.md | 2 +- .../cosmos3/generator/audiovisual/README.md | 4 ++-- .../audiovisual/run_with_trt_llm.ipynb | 14 +++++++------- 4 files changed, 20 insertions(+), 19 deletions(-) diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 137d7945..c5ffde15 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -177,7 +177,8 @@ in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155), Transfer in [#16394](https://github.com/NVIDIA/TensorRT-LLM/pull/16394), and Action in [#17325](https://github.com/NVIDIA/TensorRT-LLM/pull/17325). -Use a TensorRT-LLM checkout or package that includes those changes. +These changes are merged on TensorRT-LLM `main`; use a current checkout or a +package that includes them. Install TensorRT-LLM following its upstream documentation. @@ -188,7 +189,7 @@ Cosmos3 VisualGen change before it is available in your installed package or release image. ```bash -apt-get update && apt-get -y install ffmpeg git git-lfs +apt-get update && apt-get -y install git git-lfs git lfs install git clone https://github.com/NVIDIA/TensorRT-LLM.git @@ -207,6 +208,7 @@ docker run --rm -it \ nvcr.io/nvidia/tensorrt-llm/devel: # Inside the container: +apt-get update && apt-get -y install ffmpeg python3 scripts/build_wheel.py --use_ccache --skip_building_wheel --linking_install_binary pip install -e . ``` @@ -229,7 +231,7 @@ pip install opencv-python-headless==5.0.0.93 Set the TensorRT-LLM source root for the shared VisualGen config YAMLs: ```bash -export TRTLLM_ROOT="${TRTLLM_ROOT:-$PWD/TensorRT-LLM}" +export TRTLLM_ROOT="${TRTLLM_ROOT:-$PWD}" export COSMOS3_TRTLLM_PORT="${COSMOS3_TRTLLM_PORT:-8000}" ``` @@ -268,12 +270,11 @@ trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous `/v1/videos`, and `/v1/images/generations`. The older `/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. -The audiovisual notebook uses the validated video -generation endpoint for text-to-image, text-to-video, image-to-video, -video-to-video, and synchronized audio. Cosmos3 text-to-image is sent as a -one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; the notebook -sends it as `num_frames=1`, `seconds=1`, and `fps=8` to satisfy the video request -schema while preserving a single generated frame. Image-to-video and +The audiovisual notebook uses `/v1/videos/sync` for text-to-image, text-to-video, +image-to-video, video-to-video, and synchronized audio. Cosmos3 text-to-image is +sent as a one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; +the notebook sends it as `num_frames=1`, `seconds=1`, and `fps=8` to satisfy the +video request schema while preserving a single generated frame. Image-to-video and video-to-video upload their reference media as multipart `input_reference`; TensorRT-LLM classifies the reference by content. Synchronized audio is enabled with `enable_audio: true` in `extra_params` and is muxed into the output video. diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 059a1d0d..75126ac9 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -238,7 +238,7 @@ older `/v1/videos/generations` spelling is a deprecated alias. The asynchronous dynamics from the checked-in AV observation clip. Both notebooks decode TensorRT-LLM's `safetensors` response, write the rollout -video, and inspects the named `action` tensor. This response contract is not the +video, and inspect the named `action` tensor. This response contract is not the top-level action JSON used by vLLM-Omni. ## Run with vLLM-Omni diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index e1d23012..4be2b078 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -272,7 +272,7 @@ prompt = json.load(open("assets/prompts/text2video/robot_kitchen.json")) negative = json.load(open("assets/negative_prompts/text2video/neg_prompt.json")) response = requests.post( - "http://localhost:8000/v1/videos/generations", + "http://localhost:8000/v1/videos/sync", json={ "prompt": json.dumps(prompt, ensure_ascii=True, separators=(",", ":")), "negative_prompt": json.dumps(negative, ensure_ascii=True, separators=(",", ":")), @@ -319,7 +319,7 @@ v2v_negative = json.load(open("assets/negative_prompts/image2video/neg_prompt.js with source_video.open("rb") as video_file: response = requests.post( - "http://localhost:8000/v1/videos/generations", + "http://localhost:8000/v1/videos/sync", data={ "prompt": json.dumps(v2v_prompt, ensure_ascii=True, separators=(",", ":")), "negative_prompt": json.dumps(v2v_negative, ensure_ascii=True, separators=(",", ":")), diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index 5c5ca042..d18e5ab0 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -27,7 +27,7 @@ "source": [ "## 1. Prerequisites\n", "\n", - "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/generations`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`. Image-to-video and video-to-video send their reference media as multipart `input_reference`.\n", + "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/sync`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`. Image-to-video and video-to-video send their reference media as multipart `input_reference`.\n", "\n", "Install the `ffmpeg` CLI in the server environment. TensorRT-LLM uses it to encode MP4 and mux synchronized audio; without it, the fallback AVI encoder drops audio.\n", "\n", @@ -52,10 +52,10 @@ "\n", "### Cosmos3-Nano\n", "\n", - "From the repository root, with TensorRT-LLM installed in the active environment:\n", + "From the TensorRT-LLM repository root, with TensorRT-LLM installed in the active environment:\n", "\n", "```bash\n", - "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD/TensorRT-LLM}\"\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", "\n", "trtllm-serve nvidia/Cosmos3-Nano \\\n", " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", @@ -67,7 +67,7 @@ "`Cosmos3-Super` uses the four-GPU config from TensorRT-LLM. The config sets `cfg_size=2`, `ulysses_size=2`, and `parallel_vae_size=4`, so launch exactly four processes.\n", "\n", "```bash\n", - "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD/TensorRT-LLM}\"\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", "\n", "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \\\n", " nvidia/Cosmos3-Super \\\n", @@ -188,7 +188,7 @@ "\n", "\n", "def video_api_url(base_url: str) -> str:\n", - " return f\"{api_root_url(base_url)}/videos/generations\"\n", + " return f\"{api_root_url(base_url)}/videos/sync\"\n", "\n", "\n", "for model, base_url in TRTLLM_ENDPOINTS.items():\n", @@ -196,7 +196,7 @@ " print(model)\n", " print(\" api root:\", api_root_url(base_url))\n", " print(\" health:\", health_url(base_url))\n", - " print(\" videos generations:\", video_api_url(base_url))\n", + " print(\" videos sync:\", video_api_url(base_url))\n", " print(\" scheme:\", parsed.scheme)\n", " print(\" host:\", parsed.netloc)\n" ], @@ -297,7 +297,7 @@ "\n", "\n", "def video_api_url(base_url: str) -> str:\n", - " return f\"{api_root_url(base_url)}/videos/generations\"\n", + " return f\"{api_root_url(base_url)}/videos/sync\"\n", "\n", "\n", "IMAGE_EXTENSIONS = {\".jpg\", \".jpeg\", \".png\", \".webp\", \".bmp\"}\n", From 394af6c446b5fb9d3a2d5d420ef6bad0dcb56c24 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Mon, 7 Sep 2026 13:33:32 -0700 Subject: [PATCH 08/14] docs: validate TensorRT-LLM Cosmos3 recipes --- cookbooks/cosmos3/README.md | 31 ++-- cookbooks/cosmos3/generator/action/README.md | 34 +++- .../action/run_fd_with_trt_llm.ipynb | 30 +++- .../action/run_id_with_trt_llm.ipynb | 22 ++- .../cosmos3/generator/audiovisual/README.md | 48 +++-- .../audiovisual/run_with_trt_llm.ipynb | 170 +++++++++++++----- .../cosmos3/generator/transfer/README.md | 20 ++- .../run_video_transfer_with_trt_llm.ipynb | 16 +- 8 files changed, 266 insertions(+), 105 deletions(-) diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index c5ffde15..a1b4e58a 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -252,7 +252,7 @@ torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \ --port "$COSMOS3_TRTLLM_PORT" ``` -**Four-step distilled T2I** (single GPU; 1024×1024, one-frame warmup): +**Four-step distilled T2I** (single GPU; 1024×1024 image warmup): ```bash trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \ @@ -264,37 +264,40 @@ trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \ ```bash trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ + --enable_visual_gen \ --port "$COSMOS3_TRTLLM_PORT" ``` The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous `/v1/videos`, and `/v1/images/generations`. The older `/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. -The audiovisual notebook uses `/v1/videos/sync` for text-to-image, text-to-video, -image-to-video, video-to-video, and synchronized audio. Cosmos3 text-to-image is -sent as a one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; -the notebook sends it as `num_frames=1`, `seconds=1`, and `fps=8` to satisfy the -video request schema while preserving a single generated frame. Image-to-video and -video-to-video upload their reference media as multipart `input_reference`; -TensorRT-LLM classifies the reference by content. Synchronized audio is enabled +The audiovisual notebook uses `/v1/images/generations` for text-to-image and +`/v1/videos/sync` for text-to-video, image-to-video, video-to-video, and +synchronized audio. Text-to-image sets `extra_params.output_type="image"` and +returns a base64-encoded PNG. Image-to-video uploads multipart +`image_reference`; video-to-video uses `video_reference`. +Synchronized audio is enabled with `enable_audio: true` in `extra_params` and is muxed into the output video. -Keep `ffmpeg` on the server `PATH`: without it, TensorRT-LLM falls back to a -video-only AVI encoder and cannot preserve generated audio. Requests send +Every video request explicitly selects MP4. Keep `ffmpeg` on the server `PATH`: +without it, the request fails early instead of returning browser-incompatible +AVI or dropping generated audio. Requests send Cosmos3 controls through `extra_params`, so use a TensorRT-LLM build that includes the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. Transfer uses the synchronous `/v1/videos/sync` route. For server-derived edge -or blur, upload the raw source video as multipart `input_reference` and set the +or blur, upload the raw source video as multipart `video_reference` and set the corresponding `extra_params` hint to `true`. For a precomputed edge, blur, depth, segmentation, or WSM control, base64-encode the control inside its hint; -no `input_reference` is needed. The server decodes inline media to bytes at the +no top-level reference is needed. The server decodes inline media to bytes at the HTTP boundary. TensorRT-LLM uses `use_guardrails` for its per-request safety switch; `guardrails`, `control_path`, and other vLLM-Omni-only names are not interchangeable. -Action requests use the same synchronous route and upload an image or video as -`input_reference`. Because an action trajectory cannot be represented in MP4 or +Action requests use the same synchronous route and upload an image as +`image_reference` or a video as `video_reference`. Use a structured action +caption in `prompt`; current TensorRT-LLM ignores the legacy `view_point` field. +Because an action trajectory cannot be represented in MP4 or AVI, `format=auto` resolves to `safetensors`; the payload contains named `video`, `action`, and `frame_rate` tensors. The asynchronous `/v1/videos` route also supports this payload: poll `GET /v1/videos/{id}`, then download it from diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 75126ac9..44878b0a 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -175,8 +175,10 @@ have available. Set up and launch the VisualGen server with the one-GPU Nano config: [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). Action requests -upload the conditioning image or video as multipart `input_reference` and put -the mode-specific fields under `extra_params`. +upload a conditioning image as multipart `image_reference` or a video as +`video_reference`, and put the mode-specific fields under `extra_params`. +Use the structured prompt contract shown below; the legacy `view_point` field +is ignored by current TensorRT-LLM. ```python import json @@ -188,26 +190,43 @@ from safetensors.torch import load as load_safetensors action_root = Path("cookbooks/cosmos3/generator/action") image_path = action_root / "assets/images/av_0.jpg" actions = json.loads((action_root / "assets/actions/av_traj_forward.json").read_text()) +prompt = json.dumps( + { + "cinematography": { + "framing": "This video is captured from a first-person perspective looking at the scene." + }, + "actions": [ + { + "time": "0:00-0:06", + "description": "The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.", + } + ], + "duration": "6s", + "fps": 10.0, + "resolution": {"H": 480, "W": 832}, + "aspect_ratio": "16,9", + }, + separators=(",", ":"), +) with image_path.open("rb") as image_file: response = requests.post( "http://localhost:8000/v1/videos/sync", data={ - "prompt": "You are an autonomous vehicle planning system.", + "prompt": prompt, "format": "safetensors", "response_format": "file", - "seed": "0", + "seed": "8", "extra_params": json.dumps( { "action_mode": "forward_dynamics", "domain_name": "av", "action": actions, - "view_point": "ego_view", "use_guardrails": True, } ), }, - files={"input_reference": (image_path.name, image_file, "image/jpeg")}, + files={"image_reference": (image_path.name, image_file, "image/jpeg")}, headers={"Accept": "application/octet-stream"}, ) response.raise_for_status() @@ -215,6 +234,9 @@ payload = load_safetensors(response.content) print(payload["video"].shape, payload["action"].shape, payload["frame_rate"].item()) ``` +The checked-in AV example uses seed 8, which was substantially more coherent +than seed 0 in validation; seed 0 degraded badly near the end of the rollout. + The `av` domain preset supplies its 60-step, 9D, 480p/10 fps recipe; other recognized domains similarly fill omitted action width, chunk, resolution, and frame rate. Inverse dynamics predicts its action tensor, while forward dynamics diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb index 6c8d7dbc..d5839e01 100644 --- a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -135,7 +135,9 @@ " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", "\n", "\n", - "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", + "def submit_action(\n", + " prompt: str, reference_path: Path, extra_params: dict, output_path: Path, *, seed: int = 0\n", + ") -> dict:\n", " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", " reference_path = reference_path.resolve()\n", " if not reference_path.exists():\n", @@ -146,7 +148,7 @@ " \"prompt\": prompt,\n", " \"format\": \"safetensors\",\n", " \"response_format\": \"file\",\n", - " \"seed\": \"0\",\n", + " \"seed\": str(seed),\n", " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", " }\n", " headers = {\"Accept\": \"application/octet-stream\"}\n", @@ -156,7 +158,7 @@ " response = requests.post(\n", " video_api_url(),\n", " data=form,\n", - " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", + " files={\"image_reference\": (reference_path.name, reference_file, media_type)},\n", " headers=headers,\n", " timeout=3600,\n", " )\n", @@ -210,12 +212,28 @@ "actions = json.loads(action_path.read_text())\n", "assert len(actions) == 60 and all(len(row) == 9 for row in actions)\n", "\n", - "prompt = \"You are an autonomous vehicle planning system.\"\n", + "prompt = json.dumps(\n", + " {\n", + " \"cinematography\": {\n", + " \"framing\": \"This video is captured from a first-person perspective looking at the scene.\"\n", + " },\n", + " \"actions\": [\n", + " {\n", + " \"time\": \"0:00-0:06\",\n", + " \"description\": \"The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.\",\n", + " }\n", + " ],\n", + " \"duration\": \"6s\",\n", + " \"fps\": 10.0,\n", + " \"resolution\": {\"H\": 480, \"W\": 832},\n", + " \"aspect_ratio\": \"16,9\",\n", + " },\n", + " separators=(\",\", \":\"),\n", + ")\n", "extra_params = {\n", " \"action_mode\": \"forward_dynamics\",\n", " \"domain_name\": \"av\",\n", " \"action\": actions,\n", - " \"view_point\": \"ego_view\",\n", " \"use_guardrails\": True,\n", "}\n", "print(\"image:\", image_path)\n", @@ -239,11 +257,13 @@ "outputs": [], "source": [ "wait_for_server()\n", + "# Seed 8 is substantially more coherent than seed 0 for this checked-in AV rollout.\n", "fd_payload = submit_action(\n", " prompt,\n", " image_path,\n", " extra_params,\n", " OUTPUT_ROOT / \"forward_dynamics_av.safetensors\",\n", + " seed=8,\n", ")\n" ] }, diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb index 2bf3cd29..9bd6c231 100644 --- a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -155,7 +155,7 @@ " response = requests.post(\n", " video_api_url(),\n", " data=form,\n", - " files={\"input_reference\": (reference_path.name, reference_file, media_type)},\n", + " files={\"video_reference\": (reference_path.name, reference_file, media_type)},\n", " headers=headers,\n", " timeout=3600,\n", " )\n", @@ -208,11 +208,27 @@ "if not video_path.exists():\n", " raise FileNotFoundError(video_path)\n", "\n", - "prompt = \"You are an autonomous vehicle planning system.\"\n", + "prompt = json.dumps(\n", + " {\n", + " \"cinematography\": {\n", + " \"framing\": \"This video is captured from a first-person perspective looking at the scene.\"\n", + " },\n", + " \"actions\": [\n", + " {\n", + " \"time\": \"0:00-0:06\",\n", + " \"description\": \"The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.\",\n", + " }\n", + " ],\n", + " \"duration\": \"6s\",\n", + " \"fps\": 10.0,\n", + " \"resolution\": {\"H\": 480, \"W\": 832},\n", + " \"aspect_ratio\": \"16,9\",\n", + " },\n", + " separators=(\",\", \":\"),\n", + ")\n", "extra_params = {\n", " \"action_mode\": \"inverse_dynamics\",\n", " \"domain_name\": \"av\",\n", - " \"view_point\": \"ego_view\",\n", " \"use_guardrails\": True,\n", "}\n", "print(\"observation:\", video_path)\n", diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index 4be2b078..9399c6a2 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -284,6 +284,8 @@ response = requests.post( "guidance_scale": 6.0, "max_sequence_length": 4096, "seed": 0, + "format": "mp4", + "response_format": "file", "extra_params": { "use_resolution_template": False, "use_duration_template": False, @@ -291,20 +293,26 @@ response = requests.post( "use_guardrails": True, }, }, + headers={"Accept": "video/mp4"}, ) response.raise_for_status() -suffix = ".avi" if "x-msvideo" in response.headers.get("content-type", "") else ".mp4" -Path(f"/tmp/cosmos3_t2v_trtllm{suffix}").write_bytes(response.content) +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_t2v_trtllm.mp4").write_bytes(response.content) ``` For image-to-video, post multipart form data to the same endpoint with the -reference image under `input_reference`. To generate synchronized audio for a +reference image under `image_reference`. To generate synchronized audio for a text-to-video or image-to-video request, add `"enable_audio": True` to `extra_params`. Keep `ffmpeg` installed in the server environment so TensorRT-LLM -can mux the generated audio into the MP4; its fallback AVI encoder is video-only. +can mux the generated audio into MP4. Explicit `format=mp4` makes a missing +encoder fail early instead of returning browser-incompatible AVI. -For video-to-video, upload an MP4 or AVI reference under `input_reference`. -TensorRT-LLM classifies the upload by content and forwards the encoded bytes to +For video-to-video, upload an MP4 reference under `video_reference`. +TensorRT-LLM forwards the encoded bytes to the Cosmos3 workers, which decode the conditioning window with NVDEC: ```python @@ -330,6 +338,8 @@ with source_video.open("rb") as video_file: "guidance_scale": "6.0", "max_sequence_length": "4096", "seed": "0", + "format": "mp4", + "response_format": "file", "extra_params": json.dumps( { "use_resolution_template": False, @@ -342,12 +352,16 @@ with source_video.open("rb") as video_file: separators=(",", ":"), ), }, - files={"input_reference": (source_video.name, video_file, "video/mp4")}, - headers={"Accept": "video/mp4, video/x-msvideo"}, + files={"video_reference": (source_video.name, video_file, "video/mp4")}, + headers={"Accept": "video/mp4"}, ) response.raise_for_status() -suffix = ".avi" if "x-msvideo" in response.headers.get("content-type", "") else ".mp4" -Path(f"/tmp/cosmos3_v2v_trtllm{suffix}").write_bytes(response.content) +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_v2v_trtllm.mp4").write_bytes(response.content) ``` `condition_video_latent_indexes` identifies clean latent frames in the output; @@ -355,15 +369,17 @@ with `[0, 1]`, TensorRT-LLM consumes the first five pixel frames from the input. Set `condition_video_keep` to `"last"` to condition on the corresponding tail window instead. -For text-to-image, use the same video generation endpoint with `num_frames=1`, -`seconds=1`, and `fps=8`; TensorRT-LLM Cosmos3 returns a one-frame video -response for this path. `num_frames` is passed explicitly so the server does not -derive an eight-frame clip from `seconds * fps`. +For text-to-image, send JSON to `/v1/images/generations` with `format=png`, +`response_format=b64_json`, and `extra_params.output_type="image"`, then +base64-decode `data[0].b64_json`. This produces a PNG image instead of wrapping +one frame in a video container. For `nvidia/Cosmos3-Super-Text2Image-4Step`, use the one-GPU -`cosmos3-t2i-1gpu.yaml` config and a 1024×1024 one-frame request. For +`cosmos3-t2i-1gpu.yaml` config and a 1024×1024 image request. For `nvidia/Cosmos3-Super-Image2Video-4Step`, use the checkpoint's default -1280×720, 189-frame, 24 fps deployment shape. Both checkpoints own a fixed +1280×720, 189-frame, 24 fps deployment shape. Launch distilled I2V with +`--enable_visual_gen`; otherwise that omni checkpoint can enter the LLM/MPI +executor. Both checkpoints own a fixed four-step schedule with guidance baked into the weights, so omit `num_inference_steps` and `guidance_scale`. Leave `use_system_prompt` unset for distilled I2V so TensorRT-LLM applies the checkpoint-declared default. diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index d18e5ab0..741f14d4 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -27,9 +27,9 @@ "source": [ "## 1. Prerequisites\n", "\n", - "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/sync`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`. Image-to-video and video-to-video send their reference media as multipart `input_reference`.\n", + "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Text-to-image uses `/v1/images/generations`; video generation uses `/v1/videos/sync`. Image-to-video uploads multipart `image_reference`, while video-to-video uploads multipart `video_reference`.\n", "\n", - "Install the `ffmpeg` CLI in the server environment. TensorRT-LLM uses it to encode MP4 and mux synchronized audio; without it, the fallback AVI encoder drops audio.\n", + "Install the `ffmpeg` CLI in the server environment. This notebook explicitly requests browser-compatible MP4, so TensorRT-LLM returns an actionable error if the encoder is unavailable; `ffmpeg` also muxes synchronized audio.\n", "\n", "Generator requires the Guardrail. Request access to the gated [nvidia/Cosmos-1.0-Guardrail](https://huggingface.co/nvidia/Cosmos-1.0-Guardrail) HF repository before running these examples. TensorRT-LLM loads guardrails by default; to disable them, set `TRTLLM_DISABLE_COSMOS3_GUARDRAILS=1` before starting the server or set `use_guardrails` to `False` in the request `extra_params`.\n", "\n", @@ -80,7 +80,7 @@ "\n", "### Four-step distilled text-to-image\n", "\n", - "Use the dedicated one-GPU T2I config so warmup compiles the 1024×1024, one-frame shape:\n", + "Use the dedicated one-GPU T2I config so warmup compiles the 1024×1024 image shape:\n", "\n", "```bash\n", "trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \\\n", @@ -90,10 +90,10 @@ "\n", "### Four-step distilled image-to-video\n", "\n", - "The published 720p × 189-frame I2V student uses the default one-GPU deployment:\n", + "The published 720p × 189-frame I2V student uses one GPU. Pass `--enable_visual_gen` so this omni checkpoint is routed to VisualGen instead of the LLM/MPI executor:\n", "\n", "```bash\n", - "trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step --port 8000\n", + "trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step --enable_visual_gen --port 8000\n", "```\n" ], "id": "start-server" @@ -191,11 +191,16 @@ " return f\"{api_root_url(base_url)}/videos/sync\"\n", "\n", "\n", + "def image_api_url(base_url: str) -> str:\n", + " return f\"{api_root_url(base_url)}/images/generations\"\n", + "\n", + "\n", "for model, base_url in TRTLLM_ENDPOINTS.items():\n", " parsed = urlparse(api_root_url(base_url))\n", " print(model)\n", " print(\" api root:\", api_root_url(base_url))\n", " print(\" health:\", health_url(base_url))\n", + " print(\" images:\", image_api_url(base_url))\n", " print(\" videos sync:\", video_api_url(base_url))\n", " print(\" scheme:\", parsed.scheme)\n", " print(\" host:\", parsed.netloc)\n" @@ -319,8 +324,8 @@ " \"use_system_prompt\": False,\n", " \"use_guardrails\": True,\n", "}\n", - "# TensorRT-LLM Cosmos3 uses the stable video endpoint path here. Text-to-image\n", - "# generation is represented as a one-frame video request.\n", + "# TensorRT-LLM uses the image endpoint for T2I and the synchronous video\n", + "# endpoint for T2V, I2V, and V2V.\n", "ASSET_SETS = {\n", " \"t2i_nano\": {\n", " \"model\": \"Cosmos3-Nano\",\n", @@ -447,6 +452,8 @@ " extra_params = dict(TRTLLM_EXTRA_PARAMS)\n", " extra_params[\"enable_audio\"] = bool(spec.get(\"enable_audio\", False))\n", " extra_params.update(spec.get(\"extra_params\", {}))\n", + " if spec[\"mode\"] == \"text2image\":\n", + " extra_params[\"output_type\"] = \"image\"\n", " payload = {\n", " \"model_mode\": spec[\"mode\"],\n", " \"name\": use_case,\n", @@ -455,9 +462,8 @@ " **FIXED_SAMPLING,\n", " }\n", " if spec[\"mode\"] == \"text2image\":\n", - " payload[\"num_frames\"] = 1\n", - " payload[\"fps\"] = 8\n", - " payload[\"seconds\"] = 1\n", + " payload.pop(\"num_frames\")\n", + " payload.pop(\"fps\")\n", " else:\n", " negative_prompt_path = asset_path(\n", " spec.get(\"negative_prompt\") or f\"assets/negative_prompts/{spec['mode']}/neg_prompt.json\"\n", @@ -480,7 +486,7 @@ " print(f\"output: {output_dir}\")\n", " print(f\"prompt: {prompt_path.relative_to(COSMOS_ROOT)}\")\n", " if spec[\"mode\"] == \"text2image\":\n", - " print(\"note: TensorRT-LLM Cosmos3 text-to-image is served as a one-frame video response\")\n", + " print(\"note: TensorRT-LLM Cosmos3 text-to-image uses /v1/images/generations\")\n", " if \"vision_path\" in payload:\n", " image_display_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", " print(f\"image: {image_display_path.relative_to(COSMOS_ROOT)}\")\n", @@ -491,7 +497,10 @@ " print(\"note: condition_video_latent_indexes=[0, 1] consumes five input pixel frames\")\n", " if payload[\"extra_params\"][\"enable_audio\"]:\n", " print(\"audio: enabled; ffmpeg is required on the server to mux it into MP4\")\n", - " preview_keys = [\"model_mode\", \"name\", \"num_steps\", \"guidance\", \"fps\", \"num_frames\", \"max_sequence_length\", \"resolution\", \"aspect_ratio\", \"seed\", \"extra_params\"]\n", + " preview_keys = [\"model_mode\", \"name\", \"num_steps\", \"guidance\"]\n", + " if spec[\"mode\"] != \"text2image\":\n", + " preview_keys.extend([\"fps\", \"num_frames\"])\n", + " preview_keys.extend([\"max_sequence_length\", \"resolution\", \"aspect_ratio\", \"seed\", \"extra_params\"])\n", " if \"seconds\" in payload:\n", " preview_keys.insert(5, \"seconds\")\n", " print(json.dumps({k: payload[k] for k in preview_keys}, indent=2))\n", @@ -527,6 +536,8 @@ " \"num_frames\": payload[\"num_frames\"],\n", " \"max_sequence_length\": payload[\"max_sequence_length\"],\n", " \"seed\": payload[\"seed\"],\n", + " \"format\": \"mp4\",\n", + " \"response_format\": \"file\",\n", " }\n", " # Distilled checkpoints own their fixed step count and guidance. Omission is\n", " # intentional: conflicting request values are rejected by TensorRT-LLM.\n", @@ -540,21 +551,37 @@ " return body\n", "\n", "\n", + "def build_trtllm_image_body(payload: dict) -> dict:\n", + " height, width = payload_dimensions(payload)\n", + " body = {\n", + " \"prompt\": payload[\"prompt\"],\n", + " \"size\": f\"{width}x{height}\",\n", + " \"format\": \"png\",\n", + " \"response_format\": \"b64_json\",\n", + " \"max_sequence_length\": payload[\"max_sequence_length\"],\n", + " \"seed\": payload[\"seed\"],\n", + " \"extra_params\": payload[\"extra_params\"],\n", + " }\n", + " if payload.get(\"num_steps\") is not None:\n", + " body[\"num_inference_steps\"] = payload[\"num_steps\"]\n", + " if payload.get(\"guidance\") is not None:\n", + " body[\"guidance_scale\"] = payload[\"guidance\"]\n", + " if payload.get(\"negative_prompt\") is not None:\n", + " body[\"negative_prompt\"] = payload[\"negative_prompt\"]\n", + " return body\n", + "\n", + "\n", "def _auth_headers() -> list[str]:\n", " api_key = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"\")\n", " return [\"-H\", f\"Authorization: Bearer {api_key}\"] if api_key else []\n", "\n", "\n", - "def _video_extension_from_response(header_text: str, content_path: Path) -> str:\n", - " lowered = header_text.lower()\n", - " if \"video/x-msvideo\" in lowered or \"video/avi\" in lowered:\n", - " return \".avi\"\n", + "def _require_mp4_response(header_text: str, content_path: Path) -> None:\n", + " if \"video/mp4\" not in header_text.lower():\n", + " raise RuntimeError(f\"Expected video/mp4 response, got headers:\\n{header_text}\")\n", " head = content_path.read_bytes()[:16]\n", - " if head.startswith(b\"RIFF\") and b\"AVI\" in head[:12]:\n", - " return \".avi\"\n", - " if len(head) >= 12 and head[4:8] == b\"ftyp\":\n", - " return \".mp4\"\n", - " return \".mp4\"\n", + " if len(head) < 12 or head[4:8] != b\"ftyp\":\n", + " raise RuntimeError(\"TensorRT-LLM returned a non-MP4 payload\")\n", "\n", "\n", "def post_video(*, payload_path: Path, payload: dict, output_stem: Path, model: str) -> Path:\n", @@ -577,7 +604,7 @@ " \"-D\",\n", " str(header_path),\n", " \"-H\",\n", - " \"Accept: video/mp4, video/x-msvideo, application/octet-stream\",\n", + " \"Accept: video/mp4\",\n", " ]\n", " cmd += _auth_headers()\n", "\n", @@ -590,10 +617,12 @@ " if payload[\"model_mode\"] == \"image2video\":\n", " reference_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", " reference_type = \"image/jpeg\"\n", + " reference_field = \"image_reference\"\n", " else:\n", " reference_path = resolve_payload_path(payload_path, payload[\"video_path\"])\n", " reference_type = \"video/mp4\"\n", - " cmd += [\"-F\", f\"input_reference=@{reference_path};type={reference_type}\"]\n", + " reference_field = \"video_reference\"\n", + " cmd += [\"-F\", f\"{reference_field}=@{reference_path};type={reference_type}\"]\n", " else:\n", " cmd += [\"-H\", \"Content-Type: application/json\"]\n", " cmd += [\"-d\", json.dumps(body, separators=(\",\", \":\"))]\n", @@ -605,21 +634,66 @@ " raise RuntimeError(f\"TensorRT-LLM request failed with exit code {result.returncode}; see {error_path}\")\n", "\n", " header_text = header_path.read_text() if header_path.exists() else \"\"\n", - " ext = _video_extension_from_response(header_text, tmp_path)\n", - " output_path = output_stem.with_suffix(ext)\n", + " _require_mp4_response(header_text, tmp_path)\n", + " output_path = output_stem.with_suffix(\".mp4\")\n", " if output_path.exists():\n", " output_path.unlink()\n", " tmp_path.replace(output_path)\n", " return output_path\n", "\n", "\n", + "def post_image(*, payload: dict, output_stem: Path, model: str) -> Path:\n", + " url = image_api_url(TRTLLM_ENDPOINTS[model])\n", + " response_path = Path(f\"{output_stem}.response.json\")\n", + " error_path = Path(f\"{output_stem}.error.txt\")\n", + " for path in [response_path, error_path]:\n", + " if path.exists():\n", + " path.unlink()\n", + "\n", + " cmd = [\n", + " \"curl\",\n", + " \"-sS\",\n", + " \"--fail-with-body\",\n", + " \"-X\",\n", + " \"POST\",\n", + " url,\n", + " \"-H\",\n", + " \"Content-Type: application/json\",\n", + " \"-d\",\n", + " json.dumps(build_trtllm_image_body(payload), separators=(\",\", \":\")),\n", + " \"-o\",\n", + " str(response_path),\n", + " ]\n", + " cmd += _auth_headers()\n", + " result = subprocess.run(cmd, text=True, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n", + " if result.returncode != 0:\n", + " server_body = response_path.read_text() if response_path.exists() else \"\"\n", + " error_path.write_text(server_body + (result.stdout or \"\") + (result.stderr or \"\"))\n", + " raise RuntimeError(f\"TensorRT-LLM request failed with exit code {result.returncode}; see {error_path}\")\n", + "\n", + " response = json.loads(response_path.read_text())\n", + " try:\n", + " image_bytes = base64.b64decode(response[\"data\"][0][\"b64_json\"], validate=True)\n", + " except (KeyError, IndexError, TypeError, ValueError) as exc:\n", + " raise RuntimeError(f\"Unexpected TensorRT-LLM image response: {response}\") from exc\n", + " if not image_bytes.startswith(b\"\\x89PNG\\r\\n\\x1a\\n\"):\n", + " raise RuntimeError(\"TensorRT-LLM returned a non-PNG image payload\")\n", + " output_path = output_stem.with_suffix(\".png\")\n", + " output_path.write_bytes(image_bytes)\n", + " return output_path\n", + "\n", + "\n", "def run_trtllm_payload(payload_path: Path, output_dir: str | Path, *, model: str) -> Path:\n", " payload_path = Path(payload_path)\n", " output_dir = Path(output_dir)\n", " output_dir.mkdir(parents=True, exist_ok=True)\n", " payload = json.loads(payload_path.read_text())\n", " output_stem = output_dir / payload[\"name\"]\n", - " endpoint = video_api_url(TRTLLM_ENDPOINTS[model])\n", + " endpoint = (\n", + " image_api_url(TRTLLM_ENDPOINTS[model])\n", + " if payload[\"model_mode\"] == \"text2image\"\n", + " else video_api_url(TRTLLM_ENDPOINTS[model])\n", + " )\n", " print(\"endpoint:\", endpoint)\n", " print(\"payload:\", payload_path)\n", " print(\"output stem:\", output_stem)\n", @@ -628,14 +702,16 @@ " if payload[\"model_mode\"] == \"video2video\":\n", " print(\"input video:\", resolve_payload_path(payload_path, payload[\"video_path\"]))\n", " t0 = time.time()\n", - " output_path = post_video(payload_path=payload_path, payload=payload, output_stem=output_stem, model=model)\n", + " if payload[\"model_mode\"] == \"text2image\":\n", + " output_path = post_image(payload=payload, output_stem=output_stem, model=model)\n", + " else:\n", + " output_path = post_video(payload_path=payload_path, payload=payload, output_stem=output_stem, model=model)\n", " print(f\"wrote {output_path} in {time.time() - t0:.1f}s\")\n", " return output_path\n", "\n", "\n", "def display_video(path: Path, *, width: int = 720) -> None:\n", - " suffix = path.suffix.lower()\n", - " media_type = \"video/x-msvideo\" if suffix == \".avi\" else \"video/mp4\"\n", + " media_type = \"video/mp4\"\n", " data = base64.b64encode(path.read_bytes()).decode(\"ascii\")\n", " label = html.escape(str(path))\n", " markup = f\"\"\"\n", @@ -649,15 +725,19 @@ "\n", "def view_run(output_dir: str | Path) -> None:\n", " output_dir = Path(output_dir)\n", + " images = [path for path in sorted(output_dir.rglob(\"*.png\"))]\n", " videos = [\n", " path\n", " for path in sorted(output_dir.rglob(\"*\"))\n", - " if path.suffix.lower() in {\".mp4\", \".avi\"}\n", + " if path.suffix.lower() == \".mp4\"\n", " and not path.name.endswith((\"_preview.mp4\", \"_browser.mp4\"))\n", " ]\n", - " if not videos:\n", - " print(f\"No generated videos found under {output_dir}\")\n", + " if not images and not videos:\n", + " print(f\"No generated media found under {output_dir}\")\n", " return\n", + " for src in images:\n", + " print(f\"source: {src} ({src.stat().st_size // 1024} KB)\")\n", + " display(Image(filename=str(src), width=720))\n", " for src in videos:\n", " print(f\"source: {src} ({src.stat().st_size // 1024} KB)\")\n", " display_video(src)\n" @@ -688,7 +768,7 @@ "source": [ "## Nano: Text to Image\n", "\n", - "Nano text-to-image generation using a structured JSON prompt. TensorRT-LLM Cosmos3 returns this as a one-frame video from the video generation endpoint; the request uses `num_frames=1`, `seconds=1`, and `fps=8`.\n" + "Nano text-to-image generation using a structured JSON prompt. TensorRT-LLM returns a base64-encoded PNG from `/v1/images/generations`.\n" ], "id": "nano-t2i" }, @@ -916,7 +996,7 @@ "source": [ "## Nano: Image to Video with Audio\n", "\n", - "Nano image-to-video generation using its paired coastal-road image and synchronized audio prompt. The reference image is uploaded as multipart `input_reference`, and `enable_audio` is set in `extra_params`.\n" + "Nano image-to-video generation using its paired coastal-road image and synchronized audio prompt. The reference image is uploaded as multipart `image_reference`, and `enable_audio` is set in `extra_params`.\n" ], "id": "nano-i2av" }, @@ -973,7 +1053,7 @@ "source": [ "## Nano: Video to Video\n", "\n", - "Nano video-to-video generation using a checked-in MP4 reference. TensorRT-LLM classifies the multipart `input_reference` by content and forwards the encoded video to each worker for NVDEC decoding. `condition_video_latent_indexes=[0, 1]` pins the first two output latent frames, which consumes the first five pixel frames of the reference.\n" + "Nano video-to-video generation using a checked-in MP4 `video_reference`. TensorRT-LLM forwards the encoded video to each worker for NVDEC decoding. `condition_video_latent_indexes=[0, 1]` pins the first two output latent frames, which consumes the first five pixel frames of the reference.\n" ], "id": "nano-v2v" }, @@ -1040,7 +1120,7 @@ "source": [ "## Super: Text to Image\n", "\n", - "Super text-to-image generation using the same structured JSON prompt. TensorRT-LLM Cosmos3 returns this as a one-frame video from the video generation endpoint; the request uses `num_frames=1`, `seconds=1`, and `fps=8`.\n" + "Super text-to-image generation using the same structured JSON prompt. TensorRT-LLM returns a base64-encoded PNG from `/v1/images/generations`.\n" ], "id": "super-t2i" }, @@ -1402,9 +1482,6 @@ " \"prompt\": \"assets/prompts/text2image/robot_draping.json\",\n", " \"resolution\": \"1024\",\n", " \"aspect_ratio\": \"1,1\",\n", - " \"num_frames\": 1,\n", - " \"fps\": 8,\n", - " \"seconds\": 1,\n", " },\n", " \"i2v_distilled\": {\n", " \"model\": \"Cosmos3-Super-Image2Video-4Step\",\n", @@ -1433,8 +1510,6 @@ " \"prompt\": compact_json_file(asset_path(spec[\"prompt\"])),\n", " \"resolution\": spec[\"resolution\"],\n", " \"aspect_ratio\": spec[\"aspect_ratio\"],\n", - " \"num_frames\": spec[\"num_frames\"],\n", - " \"fps\": spec[\"fps\"],\n", " \"max_sequence_length\": 4096,\n", " \"seed\": 0,\n", " \"extra_params\": {\n", @@ -1444,6 +1519,11 @@ " \"enable_audio\": False,\n", " },\n", " }\n", + " if spec[\"mode\"] == \"text2image\":\n", + " payload[\"extra_params\"][\"output_type\"] = \"image\"\n", + " if spec[\"mode\"] != \"text2image\":\n", + " payload[\"num_frames\"] = spec[\"num_frames\"]\n", + " payload[\"fps\"] = spec[\"fps\"]\n", " if \"seconds\" in spec:\n", " payload[\"seconds\"] = spec[\"seconds\"]\n", " if \"negative_prompt\" in spec:\n", @@ -1469,9 +1549,8 @@ "source": [ "## Distilled text to image\n", "\n", - "The server is compiled for 1024×1024. The stable video route returns a one-frame media response,\n", - "so the request sends `num_frames=1`, `seconds=1`, and `fps=8` while leaving the fixed sampling\n", - "fields unset.\n" + "The server is compiled for 1024×1024. The image route returns a base64-encoded PNG while\n", + "the request leaves the checkpoint's fixed sampling fields unset.\n" ] }, { @@ -1493,8 +1572,7 @@ "source": [ "check_trtllm_server(distilled_t2i_model)\n", "distilled_t2i_data = json.loads(distilled_t2i_payload.read_text())\n", - "post_video(\n", - " payload_path=distilled_t2i_payload,\n", + "post_image(\n", " payload=distilled_t2i_data,\n", " output_stem=distilled_t2i_output / \"distilled_t2i\",\n", " model=distilled_t2i_model,\n", diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index f535f400..a9f953a0 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -26,8 +26,8 @@ Environment setup is centralized in the shared Video transfer generates a target clip from a `prompt.json` caption and one or more spatial control signals. The Framework path uses `model_mode` `video2video` in a local JSON spec. TensorRT-LLM uses `POST /v1/videos/sync`. A raw source used to derive edge or -blur is multipart `input_reference`; precomputed controls are carried directly -under their `extra_params` hint and do not need an `input_reference`. +blur is multipart `video_reference`; precomputed controls are carried directly +under their `extra_params` hint and do not need a top-level reference. The vLLM-Omni path uses `POST /v1/videos/sync` and passes one or more hint keys (`edge`, `blur`, `depth`, `seg`, or `wsm`) inside `extra_params`. Cosmos Framework accepts pre-computed control videos (`control_path`) or derives active controls from a raw source video @@ -200,10 +200,10 @@ reuses the previews from [`preview_helpers.py`](./preview_helpers.py), writing o Set up and launch Nano or Super as described in the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). TensorRT-LLM -accepts a raw source video as multipart `input_reference` when edge or blur will +accepts a raw source video as multipart `video_reference` when edge or blur will be computed on the server. The checked-in assets are already precomputed controls, so this example base64-encodes the control inside its hint and does -not send an `input_reference`. The server decodes inline media at the HTTP +not send a top-level reference. The server decodes inline media at the HTTP boundary before validating and dispatching the request. ```python @@ -244,15 +244,19 @@ response = requests.post( "guidance_scale": 3.0, "max_sequence_length": 4096, "seed": 2026, - "format": "auto", + "format": "mp4", "response_format": "file", "extra_params": extra_params, }, - headers={"Accept": "video/mp4, video/x-msvideo"}, + headers={"Accept": "video/mp4"}, ) response.raise_for_status() -suffix = ".avi" if "avi" in response.headers.get("content-type", "") else ".mp4" -Path(f"/tmp/cosmos3_transfer_depth_trtllm{suffix}").write_bytes(response.content) +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_transfer_depth_trtllm.mp4").write_bytes(response.content) ``` Only edge and blur can be generated from a raw uploaded source (`"edge": true` diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb index 1b6c4048..5fe90ac7 100644 --- a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -46,9 +46,10 @@ " --port 8000\n", "```\n", "\n", - "Install `requests` in this notebook kernel if needed. The response is synchronous encoded video\n", + "Install `requests` in this notebook kernel if needed, and install `ffmpeg` in the server\n", + "environment. The response is synchronous MP4 video\n", "bytes. Every checked-in asset is already a precomputed control, so the notebook sends it once as\n", - "base64 media inside JSON `extra_params`; no `input_reference` upload is needed. TensorRT-LLM\n", + "base64 media inside JSON `extra_params`; no top-level reference upload is needed. TensorRT-LLM\n", "decodes the media back to bytes at the HTTP boundary.\n" ] }, @@ -103,7 +104,7 @@ "All checked-in controls are 16:9 and requests explicitly select `1280x720`. TensorRT-LLM can\n", "also derive the nearest supported bucket from the first precomputed control when neither dimension\n", "is set. To compute edge or blur on the server instead, upload a raw source as multipart\n", - "`input_reference` and set the corresponding hint to `true`.\n" + "`video_reference` and set the corresponding hint to `true`.\n" ] }, { @@ -232,11 +233,11 @@ " \"guidance_scale\": spec[\"guidance_scale\"],\n", " \"max_sequence_length\": 4096,\n", " \"seed\": 2026,\n", - " \"format\": \"auto\",\n", + " \"format\": \"mp4\",\n", " \"response_format\": \"file\",\n", " \"extra_params\": extra_params,\n", " }\n", - " headers = {\"Accept\": \"video/mp4, video/x-msvideo\"}\n", + " headers = {\"Accept\": \"video/mp4\"}\n", " if TRTLLM_API_KEY:\n", " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", " response = requests.post(\n", @@ -248,10 +249,11 @@ " if not response.ok:\n", " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", " content_type = response.headers.get(\"content-type\", \"\").lower()\n", - " suffix = \".avi\" if \"avi\" in content_type or response.content[:4] == b\"RIFF\" else \".mp4\"\n", " if not response.content:\n", " raise RuntimeError(\"TensorRT-LLM returned an empty media response\")\n", - " output_path = OUTPUT_ROOT / f\"transfer_{control_name}{suffix}\"\n", + " if \"video/mp4\" not in content_type or response.content[4:8] != b\"ftyp\":\n", + " raise RuntimeError(f\"Expected browser-compatible MP4, got {content_type!r}\")\n", + " output_path = OUTPUT_ROOT / f\"transfer_{control_name}.mp4\"\n", " output_path.write_bytes(response.content)\n", " print(\"control:\", control_path)\n", " print(\"extra_params:\", json.dumps({**extra_params, control_name: \"\"}, indent=2))\n", From 7ff4a8686f52e6fdd32ebf449171fec9b0695444 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Tue, 8 Sep 2026 10:23:25 -0700 Subject: [PATCH 09/14] docs: remove TensorRT-LLM distilled examples Signed-off-by: Igor Shovkun --- README.md | 2 +- cookbooks/cosmos3/README.md | 24 +- .../cosmos3/generator/audiovisual/README.md | 16 +- .../image2video/neg_prompt.json | 8 +- .../prompts/image2video/humanoid_robot.json | 4 +- .../audiovisual/run_with_trt_llm.ipynb | 212 +----------------- 6 files changed, 12 insertions(+), 254 deletions(-) diff --git a/README.md b/README.md index 555e48fe..81210d49 100644 --- a/README.md +++ b/README.md @@ -1158,7 +1158,7 @@ We are building examples that show Cosmos 3 Super/Nano/Edge capabilities end to | --- | --- | --- | --- | --- | | Generator (audiovisual) with Diffusers | Generator | Text-to-image, plus text-to-video and image-to-video each with or without synchronized sound, via `Cosmos3OmniPipeline`. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_diffusers.ipynb) | | Generator (audiovisual) with Cosmos Framework | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, through the `cosmos_framework.scripts.inference` entrypoint. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_cosmos_framework.ipynb) | -| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video and image-to-video with or without synchronized audio, video-to-video, and the published four-step T2I/I2V students against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | +| Generator (audiovisual) with TensorRT-LLM | Generator | Text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video against an OpenAI-compatible TensorRT-LLM VisualGen server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb) | | Generator (audiovisual) with vLLM-Omni | Generator | Text-to-image, text-to-video, image-to-video, and video-to-video, with supported sound modes, against an OpenAI-compatible vLLM-Omni server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_vllm_omni.ipynb) | | Generator (audiovisual) with NIM | Generator | Text2Video and Image2Video only, against the prebuilt `Cosmos3-Generator` NIM; requests use `POST /v1/infer` and decode JSON `b64_video` responses. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_nim.ipynb) | | Generator (audiovisual) with SGLang | Generator | Text-to-image, plus text-to-video and image-to-video each with sound on or off, against an OpenAI-compatible SGLang server. | [Notebook](cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | [![Render with nbviewer](https://raw.githubusercontent.com/jupyter/design/master/logos/Badges/nbviewer_badge.svg)](https://nbviewer.org/github/nvidia/cosmos/blob/main/cookbooks/cosmos3/generator/audiovisual/run_with_sglang.ipynb) | diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index a1b4e58a..7a2cc5b0 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -170,7 +170,7 @@ uv pip install --torch-backend=cu130 \ OpenAI-compatible **VisualGen** server for Generator audiovisual text-to-image, text-to-video, image-to-video, video-to-video, synchronized audio, Transfer, and -Action examples, including the published four-step T2I and I2V students. +Action examples. Initial Cosmos3 support was added in TensorRT-LLM PR [#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824), synchronized audio in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and @@ -252,22 +252,6 @@ torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \ --port "$COSMOS3_TRTLLM_PORT" ``` -**Four-step distilled T2I** (single GPU; 1024×1024 image warmup): - -```bash -trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \ - --visual_gen_args "$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-t2i-1gpu.yaml" \ - --port "$COSMOS3_TRTLLM_PORT" -``` - -**Four-step distilled I2V** (single GPU; default 1280×720, 189-frame shape): - -```bash -trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step \ - --enable_visual_gen \ - --port "$COSMOS3_TRTLLM_PORT" -``` - The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous `/v1/videos`, and `/v1/images/generations`. The older `/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. @@ -303,12 +287,6 @@ AVI, `format=auto` resolves to `safetensors`; the payload contains named `video` supports this payload: poll `GET /v1/videos/{id}`, then download it from `GET /v1/videos/{id}/content`. -The distilled checkpoints own their four-step stochastic schedules and have -classifier-free guidance baked into their weights. Leave `num_inference_steps` -and `guidance_scale` unset; conflicting values are rejected. Also leave -`use_system_prompt` unset for distilled I2V so its checkpoint-declared default -is applied. - ## TensorRT-LLM Reasoner OpenAI-compatible **reasoning** server for image and video understanding. Run diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index 9399c6a2..fe3e06de 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -374,16 +374,6 @@ For text-to-image, send JSON to `/v1/images/generations` with `format=png`, base64-decode `data[0].b64_json`. This produces a PNG image instead of wrapping one frame in a video container. -For `nvidia/Cosmos3-Super-Text2Image-4Step`, use the one-GPU -`cosmos3-t2i-1gpu.yaml` config and a 1024×1024 image request. For -`nvidia/Cosmos3-Super-Image2Video-4Step`, use the checkpoint's default -1280×720, 189-frame, 24 fps deployment shape. Launch distilled I2V with -`--enable_visual_gen`; otherwise that omni checkpoint can enter the LLM/MPI -executor. Both checkpoints own a fixed -four-step schedule with guidance baked into the weights, so omit -`num_inference_steps` and `guidance_scale`. Leave `use_system_prompt` unset for -distilled I2V so TensorRT-LLM applies the checkpoint-declared default. - The TRT-LLM notebook always sends model-specific `extra_params`, so use a TensorRT-LLM release with the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. @@ -393,10 +383,8 @@ request-level `max_sequence_length=4096` for longer structured JSON prompts. [`run_with_trt_llm.ipynb`](./run_with_trt_llm.ipynb) is the full tutorial for the TensorRT-LLM backend: it walks through text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video requests -against an already-running VisualGen server. Independent final sections cover -the published four-step T2I and I2V students without overriding their fixed -sampling recipes. Server launch options (Nano, Super, distilled T2I, and -distilled I2V) live in the +against an already-running VisualGen server. Server launch options for Nano and +Super live in the [shared environment setup guide](../../README.md#tensorrt-llm-generator). ## Run with NIM diff --git a/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json b/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json index 301387bf..2a0bd29b 100644 --- a/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json +++ b/cookbooks/cosmos3/generator/audiovisual/assets/negative_prompts/image2video/neg_prompt.json @@ -2,7 +2,7 @@ "subjects": [ { "description": "Blurry, poorly defined subjects with inconsistent shapes and unrealistic proportions.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -22,7 +22,7 @@ }, { "description": "Extremely low-quality subjects with visible rendering artifacts, broken mesh geometry, and completely unrealistic proportions throughout.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -42,7 +42,7 @@ }, { "description": "Poorly generated subjects exhibiting all hallmarks of failed neural rendering \u2014 flickering edges, inconsistent depth, and uncanny spatial relationships.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, poor color separation between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -71,7 +71,7 @@ "aesthetics": { "composition": "Cluttered, poorly framed composition with no clear focal point. Important elements are cut off by the frame edges. The rule of thirds is completely ignored, leading to an unbalanced and visually unpleasant arrangement.", "color_scheme": "Oversaturated, garish colors that clash violently. Color banding is visible in gradient areas. The overall palette feels artificial and digitally processed rather than natural.", - "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels flat and sterile despite attempting to portray dynamic action.", + "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels lifeless and sterile despite attempting to portray dynamic action.", "patterns": "Visible tiling artifacts in textures, moir\u00e9 patterns, and aliasing on edges." }, "cinematography": { diff --git a/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json b/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json index 24f2d76b..6a14ef50 100644 --- a/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json +++ b/cookbooks/cosmos3/generator/audiovisual/assets/prompts/image2video/humanoid_robot.json @@ -92,7 +92,7 @@ } ], "transitions": [], - "temporal_caption": "The video opens with a humanoid robot standing motionless in a modern living room. At around one second, the robot begins to shift its weight and bend its knees slightly. By two seconds, it drops into a deep crouch with arms swinging back. At approximately three seconds, the robot explosively launches itself upward and backward, its white body arcing through the air as it tucks its legs tightly. Between three and five seconds, the robot completes a full backward rotation high above the living room floor, its mechanical joints visible as it spins. At five seconds, the robot's feet come back down and make contact with the rug, knees deeply bent to absorb the landing impact. From 5 to 7 seconds, the robot smoothly extends back to its full standing height, returning to a relaxed upright pose as if nothing extraordinary just occurred.", + "temporal_caption": "The video opens with a humanoid robot standing motionless in a modern living room. At around one second, the robot begins to shift its weight and bend its knees slightly. By two seconds, it drops into a deep crouch with arms swinging back. At approximately three seconds, the robot explosively launches itself upward and backward, its white body arcing through the air as it tucks its legs tightly. Between three and five seconds, the robot completes a full backward rotation high above the living room floor, its mechanical joints visible as it spins. At five seconds, the robot's feet come back down and make contact with the rug, knees deeply bent to absorb the landing impact. From five to seven seconds, the robot smoothly extends back to its full standing height, returning to a relaxed upright pose as if nothing extraordinary just occurred.", "resolution": { "W": 1280, "H": 720 @@ -100,4 +100,4 @@ "aspect_ratio": "16,9", "duration": "7s", "fps": 24 - } + } \ No newline at end of file diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index 741f14d4..28d4e664 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -17,7 +17,7 @@ "\n", "This notebook calls already-running TensorRT-LLM VisualGen servers with direct `curl` requests from Python.\n", "\n", - "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video and image-to-video with or without synchronized audio, video-to-video generation, and the published four-step T2I/I2V students.\n" + "The examples are split into Cosmos3-Nano and Cosmos3-Super sections. Each section is self-contained, so you can run just one. The notebook covers TensorRT-LLM's stable server flow for text-to-image, text-to-video and image-to-video with or without synchronized audio, and video-to-video generation.\n" ], "id": "title" }, @@ -77,24 +77,7 @@ "\n", "TensorRT-LLM exposes `/health` when the server is ready. This notebook sends Cosmos3 model-specific controls through `extra_params`, so use a TensorRT-LLM release that includes the Cosmos3 VisualGen API schema.\n", "\n", - "\n", - "### Four-step distilled text-to-image\n", - "\n", - "Use the dedicated one-GPU T2I config so warmup compiles the 1024×1024 image shape:\n", - "\n", - "```bash\n", - "trtllm-serve nvidia/Cosmos3-Super-Text2Image-4Step \\\n", - " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-t2i-1gpu.yaml\" \\\n", - " --port 8000\n", - "```\n", - "\n", - "### Four-step distilled image-to-video\n", - "\n", - "The published 720p × 189-frame I2V student uses one GPU. Pass `--enable_visual_gen` so this omni checkpoint is routed to VisualGen instead of the LLM/MPI executor:\n", - "\n", - "```bash\n", - "trtllm-serve nvidia/Cosmos3-Super-Image2Video-4Step --enable_visual_gen --port 8000\n", - "```\n" + "\n" ], "id": "start-server" }, @@ -134,12 +117,6 @@ "TRTLLM_ENDPOINTS = {\n", " \"Cosmos3-Nano\": os.environ.get(\"COSMOS3_TRTLLM_NANO_BASE_URL\", DEFAULT_TRTLLM_BASE_URL),\n", " \"Cosmos3-Super\": os.environ.get(\"COSMOS3_TRTLLM_SUPER_BASE_URL\", DEFAULT_TRTLLM_BASE_URL),\n", - " \"Cosmos3-Super-Text2Image-4Step\": os.environ.get(\n", - " \"COSMOS3_TRTLLM_DISTILLED_T2I_BASE_URL\", DEFAULT_TRTLLM_BASE_URL\n", - " ),\n", - " \"Cosmos3-Super-Image2Video-4Step\": os.environ.get(\n", - " \"COSMOS3_TRTLLM_DISTILLED_I2V_BASE_URL\", DEFAULT_TRTLLM_BASE_URL\n", - " ),\n", "}\n", "\n", "os.environ[\"COSMOS3_AUDIOVISUAL_OUTPUT_ROOT\"] = str(COSMOS3_AUDIOVISUAL_OUTPUT_ROOT)\n", @@ -428,8 +405,6 @@ " return 720, 1280\n", " if payload.get(\"resolution\") == \"256\" and payload.get(\"aspect_ratio\") == \"16,9\":\n", " return 192, 320\n", - " if payload.get(\"resolution\") == \"1024\" and payload.get(\"aspect_ratio\") == \"1,1\":\n", - " return 1024, 1024\n", " raise ValueError(f\"Unsupported payload resolution/aspect ratio: {payload.get('resolution')} {payload.get('aspect_ratio')}\")\n", "\n", "\n", @@ -539,8 +514,6 @@ " \"format\": \"mp4\",\n", " \"response_format\": \"file\",\n", " }\n", - " # Distilled checkpoints own their fixed step count and guidance. Omission is\n", - " # intentional: conflicting request values are rejected by TensorRT-LLM.\n", " if payload.get(\"num_steps\") is not None:\n", " body[\"num_inference_steps\"] = payload[\"num_steps\"]\n", " if payload.get(\"guidance\") is not None:\n", @@ -1455,187 +1428,6 @@ "view_run(v2v_super_output)\n" ], "id": "super-v2v-view" - }, - { - "cell_type": "markdown", - "id": "distilled-title", - "metadata": {}, - "source": [ - "# Four-step Distilled Cosmos3-Super Examples\n", - "\n", - "These checkpoints carry a fixed four-step stochastic schedule and classifier-free guidance baked\n", - "into the weights. The request bodies deliberately omit `num_inference_steps` and\n", - "`guidance_scale`; TensorRT-LLM reads both from the checkpoint and rejects conflicting values.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-helpers", - "metadata": {}, - "outputs": [], - "source": [ - "DISTILLED_ASSET_SETS = {\n", - " \"t2i_distilled\": {\n", - " \"model\": \"Cosmos3-Super-Text2Image-4Step\",\n", - " \"mode\": \"text2image\",\n", - " \"prompt\": \"assets/prompts/text2image/robot_draping.json\",\n", - " \"resolution\": \"1024\",\n", - " \"aspect_ratio\": \"1,1\",\n", - " },\n", - " \"i2v_distilled\": {\n", - " \"model\": \"Cosmos3-Super-Image2Video-4Step\",\n", - " \"mode\": \"image2video\",\n", - " \"prompt\": \"assets/prompts/image2video/humanoid_robot.json\",\n", - " \"negative_prompt\": \"assets/negative_prompts/image2video/neg_prompt.json\",\n", - " \"image\": \"assets/images/image2video/humanoid_robot.jpg\",\n", - " \"resolution\": \"720\",\n", - " \"aspect_ratio\": \"16,9\",\n", - " \"num_frames\": 189,\n", - " \"fps\": 24,\n", - " },\n", - "}\n", - "\n", - "\n", - "def create_distilled_payload(use_case: str) -> tuple[Path, Path, str]:\n", - " spec = DISTILLED_ASSET_SETS[use_case]\n", - " payload_dir = COSMOS3_AUDIOVISUAL_OUTPUT_ROOT / \"trt_llm\" / \"payloads\" / use_case\n", - " output_dir = COSMOS3_AUDIOVISUAL_OUTPUT_ROOT / \"trt_llm\" / use_case\n", - " payload_dir.mkdir(parents=True, exist_ok=True)\n", - " output_dir.mkdir(parents=True, exist_ok=True)\n", - " payload_path = payload_dir / f\"{use_case}.json\"\n", - " payload = {\n", - " \"model_mode\": spec[\"mode\"],\n", - " \"name\": use_case,\n", - " \"prompt\": compact_json_file(asset_path(spec[\"prompt\"])),\n", - " \"resolution\": spec[\"resolution\"],\n", - " \"aspect_ratio\": spec[\"aspect_ratio\"],\n", - " \"max_sequence_length\": 4096,\n", - " \"seed\": 0,\n", - " \"extra_params\": {\n", - " \"use_resolution_template\": False,\n", - " \"use_duration_template\": False,\n", - " \"use_guardrails\": True,\n", - " \"enable_audio\": False,\n", - " },\n", - " }\n", - " if spec[\"mode\"] == \"text2image\":\n", - " payload[\"extra_params\"][\"output_type\"] = \"image\"\n", - " if spec[\"mode\"] != \"text2image\":\n", - " payload[\"num_frames\"] = spec[\"num_frames\"]\n", - " payload[\"fps\"] = spec[\"fps\"]\n", - " if \"seconds\" in spec:\n", - " payload[\"seconds\"] = spec[\"seconds\"]\n", - " if \"negative_prompt\" in spec:\n", - " payload[\"negative_prompt\"] = compact_json_file(asset_path(spec[\"negative_prompt\"]))\n", - " if \"image\" in spec:\n", - " image_path = asset_path(spec[\"image\"])\n", - " payload[\"vision_path\"] = os.path.relpath(image_path, payload_path.parent)\n", - " display(Image(filename=str(image_path), width=420))\n", - " # Do not add use_system_prompt: the I2V checkpoint declares True and the\n", - " # pipeline applies that checkpoint default only when the request leaves it unset.\n", - " payload_path.write_text(json.dumps(payload, indent=2) + \"\\n\")\n", - " print(\"model:\", spec[\"model\"])\n", - " print(\"payload:\", payload_path)\n", - " print(\"output:\", output_dir)\n", - " print(json.dumps(payload, indent=2))\n", - " return payload_path, output_dir, spec[\"model\"]\n" - ] - }, - { - "cell_type": "markdown", - "id": "distilled-t2i-title", - "metadata": {}, - "source": [ - "## Distilled text to image\n", - "\n", - "The server is compiled for 1024×1024. The image route returns a base64-encoded PNG while\n", - "the request leaves the checkpoint's fixed sampling fields unset.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-t2i-payload", - "metadata": {}, - "outputs": [], - "source": [ - "distilled_t2i_payload, distilled_t2i_output, distilled_t2i_model = create_distilled_payload(\"t2i_distilled\")\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-t2i-run", - "metadata": {}, - "outputs": [], - "source": [ - "check_trtllm_server(distilled_t2i_model)\n", - "distilled_t2i_data = json.loads(distilled_t2i_payload.read_text())\n", - "post_image(\n", - " payload=distilled_t2i_data,\n", - " output_stem=distilled_t2i_output / \"distilled_t2i\",\n", - " model=distilled_t2i_model,\n", - ")\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-t2i-view", - "metadata": {}, - "outputs": [], - "source": [ - "view_run(distilled_t2i_output)\n" - ] - }, - { - "cell_type": "markdown", - "id": "distilled-i2v-title", - "metadata": {}, - "source": [ - "## Distilled image to video\n", - "\n", - "The published deployment shape is 1280×720 with 189 frames at 24 fps. The request leaves\n", - "`use_system_prompt` unset so the checkpoint's declared `true` default is applied.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-i2v-payload", - "metadata": {}, - "outputs": [], - "source": [ - "distilled_i2v_payload, distilled_i2v_output, distilled_i2v_model = create_distilled_payload(\"i2v_distilled\")\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-i2v-run", - "metadata": {}, - "outputs": [], - "source": [ - "check_trtllm_server(distilled_i2v_model)\n", - "distilled_i2v_data = json.loads(distilled_i2v_payload.read_text())\n", - "post_video(\n", - " payload_path=distilled_i2v_payload,\n", - " payload=distilled_i2v_data,\n", - " output_stem=distilled_i2v_output / \"distilled_i2v\",\n", - " model=distilled_i2v_model,\n", - ")\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "distilled-i2v-view", - "metadata": {}, - "outputs": [], - "source": [ - "view_run(distilled_i2v_output)\n" - ] } ], "metadata": { From 7f973ea81f4b37bab2d875fa34a90313c35a9990 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Wed, 9 Sep 2026 15:20:31 -0700 Subject: [PATCH 10/14] docs: align TRT-LLM action and transfer recipes Signed-off-by: Igor Shovkun --- cookbooks/cosmos3/README.md | 5 ++-- cookbooks/cosmos3/generator/action/README.md | 28 +++---------------- .../action/run_fd_with_trt_llm.ipynb | 22 ++------------- .../action/run_id_with_trt_llm.ipynb | 19 +------------ .../cosmos3/generator/transfer/README.md | 12 +++++--- .../run_video_transfer_with_trt_llm.ipynb | 15 ++++++++-- 6 files changed, 30 insertions(+), 71 deletions(-) diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 7a2cc5b0..f0ee53e8 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -279,8 +279,9 @@ switch; `guardrails`, `control_path`, and other vLLM-Omni-only names are not interchangeable. Action requests use the same synchronous route and upload an image as -`image_reference` or a video as `video_reference`. Use a structured action -caption in `prompt`; current TensorRT-LLM ignores the legacy `view_point` field. +`image_reference` or a video as `video_reference`. For the checked-in AV +examples, use the Cosmos Framework reference prompt shown in the Action +cookbook; current TensorRT-LLM ignores the legacy `view_point` field. Because an action trajectory cannot be represented in MP4 or AVI, `format=auto` resolves to `safetensors`; the payload contains named `video`, `action`, and `frame_rate` tensors. The asynchronous `/v1/videos` route also diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 44878b0a..44af9a49 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -177,8 +177,8 @@ Set up and launch the VisualGen server with the one-GPU Nano config: [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). Action requests upload a conditioning image as multipart `image_reference` or a video as `video_reference`, and put the mode-specific fields under `extra_params`. -Use the structured prompt contract shown below; the legacy `view_point` field -is ignored by current TensorRT-LLM. +Use the same trained AV prompt as the Cosmos Framework reference; the legacy +`view_point` field is ignored by current TensorRT-LLM. ```python import json @@ -190,24 +190,7 @@ from safetensors.torch import load as load_safetensors action_root = Path("cookbooks/cosmos3/generator/action") image_path = action_root / "assets/images/av_0.jpg" actions = json.loads((action_root / "assets/actions/av_traj_forward.json").read_text()) -prompt = json.dumps( - { - "cinematography": { - "framing": "This video is captured from a first-person perspective looking at the scene." - }, - "actions": [ - { - "time": "0:00-0:06", - "description": "The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.", - } - ], - "duration": "6s", - "fps": 10.0, - "resolution": {"H": 480, "W": 832}, - "aspect_ratio": "16,9", - }, - separators=(",", ":"), -) +prompt = "You are an autonomous vehicle planning system." with image_path.open("rb") as image_file: response = requests.post( @@ -216,7 +199,7 @@ with image_path.open("rb") as image_file: "prompt": prompt, "format": "safetensors", "response_format": "file", - "seed": "8", + "seed": "0", "extra_params": json.dumps( { "action_mode": "forward_dynamics", @@ -234,9 +217,6 @@ payload = load_safetensors(response.content) print(payload["video"].shape, payload["action"].shape, payload["frame_rate"].item()) ``` -The checked-in AV example uses seed 8, which was substantially more coherent -than seed 0 in validation; seed 0 degraded badly near the end of the rollout. - The `av` domain preset supplies its 60-step, 9D, 480p/10 fps recipe; other recognized domains similarly fill omitted action width, chunk, resolution, and frame rate. Inverse dynamics predicts its action tensor, while forward dynamics diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb index d5839e01..db73988d 100644 --- a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -212,24 +212,7 @@ "actions = json.loads(action_path.read_text())\n", "assert len(actions) == 60 and all(len(row) == 9 for row in actions)\n", "\n", - "prompt = json.dumps(\n", - " {\n", - " \"cinematography\": {\n", - " \"framing\": \"This video is captured from a first-person perspective looking at the scene.\"\n", - " },\n", - " \"actions\": [\n", - " {\n", - " \"time\": \"0:00-0:06\",\n", - " \"description\": \"The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.\",\n", - " }\n", - " ],\n", - " \"duration\": \"6s\",\n", - " \"fps\": 10.0,\n", - " \"resolution\": {\"H\": 480, \"W\": 832},\n", - " \"aspect_ratio\": \"16,9\",\n", - " },\n", - " separators=(\",\", \":\"),\n", - ")\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", "extra_params = {\n", " \"action_mode\": \"forward_dynamics\",\n", " \"domain_name\": \"av\",\n", @@ -257,13 +240,12 @@ "outputs": [], "source": [ "wait_for_server()\n", - "# Seed 8 is substantially more coherent than seed 0 for this checked-in AV rollout.\n", "fd_payload = submit_action(\n", " prompt,\n", " image_path,\n", " extra_params,\n", " OUTPUT_ROOT / \"forward_dynamics_av.safetensors\",\n", - " seed=8,\n", + " seed=0,\n", ")\n" ] }, diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb index 9bd6c231..3fa07636 100644 --- a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -208,24 +208,7 @@ "if not video_path.exists():\n", " raise FileNotFoundError(video_path)\n", "\n", - "prompt = json.dumps(\n", - " {\n", - " \"cinematography\": {\n", - " \"framing\": \"This video is captured from a first-person perspective looking at the scene.\"\n", - " },\n", - " \"actions\": [\n", - " {\n", - " \"time\": \"0:00-0:06\",\n", - " \"description\": \"The vehicle drives straight forward with smooth ego motion while the road and surrounding buildings remain consistent.\",\n", - " }\n", - " ],\n", - " \"duration\": \"6s\",\n", - " \"fps\": 10.0,\n", - " \"resolution\": {\"H\": 480, \"W\": 832},\n", - " \"aspect_ratio\": \"16,9\",\n", - " },\n", - " separators=(\",\", \":\"),\n", - ")\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", "extra_params = {\n", " \"action_mode\": \"inverse_dynamics\",\n", " \"domain_name\": \"av\",\n", diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index a9f953a0..0f8cbaf8 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -222,6 +222,7 @@ extra_params = { "use_duration_template": False, "use_system_prompt": False, "use_guardrails": True, + "flow_shift": 10.0, "depth": { "control": base64.b64encode(control_path.read_bytes()).decode("ascii") }, @@ -229,6 +230,8 @@ extra_params = { "num_video_frames_per_chunk": 121, "num_conditional_frames": 1, "num_first_chunk_conditional_frames": 0, + "share_vision_temporal_positions": True, + "emphasize_control_in_prompt": False, "max_frames": 121, } @@ -240,7 +243,7 @@ response = requests.post( "size": "1280x720", "num_frames": 121, "fps": 30, - "num_inference_steps": 35, + "num_inference_steps": 50, "guidance_scale": 3.0, "max_sequence_length": 4096, "seed": 2026, @@ -265,9 +268,10 @@ control media. Precomputed edge and blur are also accepted. Use `use_guardrails`, not vLLM-Omni's `guardrails`, and send encoded control bytes rather than a server-local `control_path`. When neither output dimension is specified, TensorRT-LLM chooses the nearest supported bucket from the source or -first precomputed control's aspect ratio. The checked-in examples explicitly -request 1280×720 and their native frame count/fps; WSM uses 100 frames at 10 -fps. `/v1/videos/generations` remains only as a deprecated alias of the +first precomputed control's aspect ratio. The checked-in blur example explicitly +requests its native 4:3 bucket, 1104×832; the other examples request +1280×720. All preserve their native frame count/fps, and WSM uses 100 frames +at 10 fps. `/v1/videos/generations` remains only as a deprecated alias of the canonical blocking `/v1/videos/sync` route. ### TensorRT-LLM notebook walkthrough diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb index 5fe90ac7..7183a6ef 100644 --- a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -101,7 +101,8 @@ "| Segmentation | Base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", "| WSM | Base64 precomputed control | 100 / 10 | 1.0 / 3.0 |\n", "\n", - "All checked-in controls are 16:9 and requests explicitly select `1280x720`. TensorRT-LLM can\n", + "The blur control is 4:3 and explicitly selects `1104x832`; the other controls are 16:9 and\n", + "explicitly select `1280x720`. TensorRT-LLM can\n", "also derive the nearest supported bucket from the first precomputed control when neither dimension\n", "is set. To compute edge or blur on the server instead, upload a raw source as multipart\n", "`video_reference` and set the corresponding hint to `true`.\n" @@ -154,6 +155,7 @@ " \"fps\": 30,\n", " \"guidance_scale\": 3.0,\n", " \"control_guidance\": 1.5,\n", + " \"size\": \"1280x720\",\n", " \"hint\": {\"preset_edge_threshold\": \"medium\"},\n", " },\n", " \"blur\": {\n", @@ -163,6 +165,7 @@ " \"fps\": 30,\n", " \"guidance_scale\": 3.0,\n", " \"control_guidance\": 1.5,\n", + " \"size\": \"1104x832\",\n", " \"hint\": {\"preset_blur_strength\": \"medium\"},\n", " },\n", " \"depth\": {\n", @@ -172,6 +175,7 @@ " \"fps\": 30,\n", " \"guidance_scale\": 3.0,\n", " \"control_guidance\": 1.5,\n", + " \"size\": \"1280x720\",\n", " },\n", " \"seg\": {\n", " \"control\": \"assets/seg/control_seg.mp4\",\n", @@ -180,6 +184,7 @@ " \"fps\": 30,\n", " \"guidance_scale\": 3.0,\n", " \"control_guidance\": 2.0,\n", + " \"size\": \"1280x720\",\n", " },\n", " \"wsm\": {\n", " \"control\": \"assets/wsm/control_wsm.mp4\",\n", @@ -188,6 +193,7 @@ " \"fps\": 10,\n", " \"guidance_scale\": 1.0,\n", " \"control_guidance\": 3.0,\n", + " \"size\": \"1280x720\",\n", " },\n", "}\n", "\n", @@ -216,20 +222,23 @@ " \"use_resolution_template\": False,\n", " \"use_system_prompt\": False,\n", " \"use_guardrails\": True,\n", + " \"flow_shift\": 10.0,\n", " control_name: hint,\n", " \"control_guidance\": spec[\"control_guidance\"],\n", " \"num_video_frames_per_chunk\": spec[\"num_frames\"],\n", " \"num_conditional_frames\": 1,\n", " \"num_first_chunk_conditional_frames\": 0,\n", + " \"share_vision_temporal_positions\": True,\n", + " \"emphasize_control_in_prompt\": False,\n", " \"max_frames\": spec[\"num_frames\"],\n", " }\n", " request_body = {\n", " \"prompt\": compact_json(prompt_path),\n", " \"negative_prompt\": compact_json(negative_path),\n", - " \"size\": \"1280x720\",\n", + " \"size\": spec[\"size\"],\n", " \"num_frames\": spec[\"num_frames\"],\n", " \"fps\": spec[\"fps\"],\n", - " \"num_inference_steps\": 35,\n", + " \"num_inference_steps\": 50,\n", " \"guidance_scale\": spec[\"guidance_scale\"],\n", " \"max_sequence_length\": 4096,\n", " \"seed\": 2026,\n", From a9029df16e75267d1d6eb168832e147e102dcd29 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 18 Sep 2026 08:19:52 -0700 Subject: [PATCH 11/14] docs: reconcile TRT-LLM Action and Transfer validation Signed-off-by: Igor Shovkun --- cookbooks/cosmos3/README.md | 14 ++- cookbooks/cosmos3/generator/action/README.md | 6 + .../cosmos3/generator/transfer/README.md | 6 + .../cosmos3/generator/trtllm-validation.md | 110 ++++++++++++++++++ 4 files changed, 134 insertions(+), 2 deletions(-) create mode 100644 cookbooks/cosmos3/generator/trtllm-validation.md diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index f0ee53e8..4a27575b 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -177,8 +177,16 @@ in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155), Transfer in [#16394](https://github.com/NVIDIA/TensorRT-LLM/pull/16394), and Action in [#17325](https://github.com/NVIDIA/TensorRT-LLM/pull/17325). -These changes are merged on TensorRT-LLM `main`; use a current checkout or a -package that includes them. +These changes are merged on TensorRT-LLM `main`. The Action and Transfer +notebooks were executed against source revision +[`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e). +Use that revision to reproduce their request contract, or a newer build with +the same API. The older `799d7d42` validation used `input_reference` and does +not validate these notebooks' `image_reference` / `video_reference` fields. +See the [Action and Transfer validation record](generator/trtllm-validation.md) +for the executed cases, output dimensions, and remaining runtime limitations. +The source revision is significant: a package version of `1.3.0rc26` alone +does not establish compatibility with the action image-decoding path. Install TensorRT-LLM following its upstream documentation. @@ -194,6 +202,8 @@ git lfs install git clone https://github.com/NVIDIA/TensorRT-LLM.git cd TensorRT-LLM +# Source revision used for the Action/Transfer notebook validation below. +git checkout bca6761ab84fbcd58fc7f914eade7de48b32e35e git submodule update --init --recursive git lfs pull diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 44af9a49..63e160f3 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -180,6 +180,12 @@ upload a conditioning image as multipart `image_reference` or a video as Use the same trained AV prompt as the Cosmos Framework reference; the legacy `view_point` field is ignored by current TensorRT-LLM. +Both notebooks' current multipart fields were exercised against TensorRT-LLM +source revision `bca6761ab84fbcd58fc7f914eade7de48b32e35e`. Follow the source +pin in the shared setup; the earlier `799d7d42` run used the old +`input_reference` API. The [validation record](../trtllm-validation.md) covers +forward and inverse dynamics separately and records the guardrail limitation. + ```python import json from pathlib import Path diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index 0f8cbaf8..ed7936c3 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -274,6 +274,12 @@ requests its native 4:3 bucket, 1104×832; the other examples request at 10 fps. `/v1/videos/generations` remains only as a deprecated alias of the canonical blocking `/v1/videos/sync` route. +The [validation record](../trtllm-validation.md) reports the later execution of +all five controls at TensorRT-LLM source revision +`bca6761ab84fbcd58fc7f914eade7de48b32e35e`, including blur at 1104×832. +The earlier 1280×720 blur result used a previous request and does not describe +the checked-in notebook. + ### TensorRT-LLM notebook walkthrough [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) diff --git a/cookbooks/cosmos3/generator/trtllm-validation.md b/cookbooks/cosmos3/generator/trtllm-validation.md new file mode 100644 index 00000000..09524cb5 --- /dev/null +++ b/cookbooks/cosmos3/generator/trtllm-validation.md @@ -0,0 +1,110 @@ +# TensorRT-LLM Action and Transfer validation + +This record reconciles the saved GPU runs with the notebook requests in +[cookbook commit `7f973ea81f4b37bab2d875fa34a90313c35a9990`](https://github.com/NVIDIA/cosmos/commit/7f973ea81f4b37bab2d875fa34a90313c35a9990). +It was audited on Fri 18 Sep 2026 (Pacific) using the retained logs and +artifacts; this documentation update did not run new GPU inference. + +## Compatible runtime + +The Action and Nano Transfer notebook runs used TensorRT-LLM source revision +[`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e) +in a native Python 3.12 environment, with PyTorch `2.12.0+cu130` and +`cosmos_guardrail==0.3.0`. The runtime reported `1.3.0rc26`; that package +version alone does not identify the source revision. These were native runs, +so there is no container digest. The retained evidence does not establish a +fresh C++ rebuild or an immutable fingerprint of every installed dependency. + +At this revision, the +[request schema](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/openai_protocol.py#L2067) +and [multipart parser](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/openai_video_routes.py#L311) +accept `image_reference` and `video_reference`. +The [conversion layer](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/visual_gen_utils.py#L482) +passes uploaded media to the worker, and the +[Action image decoder](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.py#L364) +handles encoded image bytes. The earlier `799d7d42` runs used `input_reference`; +they are historical evidence, not validation of the current request format. + +The server used the checked-in `cosmos3-nano-1gpu.yaml`: + +```bash +trtllm-serve nvidia/Cosmos3-Nano \ + --visual_gen_args examples/visual_gen/configs/cosmos3-nano-1gpu.yaml \ + --host 127.0.0.1 --port 8000 +``` + +The checkpoint resolved to snapshot +`7a312c868bcce8e40b3eb40861300a9d0ba3fde1`. Each run used one H100 80GB on +`ipp2-0160`. See the [shared setup](../README.md#tensorrt-llm-generator) for +the source pin and launch commands. + +## Executed notebook requests + +The runs below executed the checked-in code cells, including request creation +and response decoding. Action used the checked-in AV assets, prompt +`You are an autonomous vehicle planning system.`, and seed 0. Transfer used +the checked-in controls and captions, seed 2026, 50 steps, and flow shift 10. +Dimensions in this table are width × height. + +| Notebook/case | Request to `POST /v1/videos/sync` | Decoded result | GPU evidence | +| --- | --- | --- | --- | +| [Forward dynamics](action/run_fd_with_trt_llm.ipynb) | Multipart `image_reference`, `format=safetensors` | `video[61,480,832,3]` uint8; `action[60,9]` float32; 10 fps | `20260909-135841-262ee2`, all 5 cells passed | +| [Inverse dynamics](action/run_id_with_trt_llm.ipynb) | Multipart `video_reference`, `format=safetensors` | `video[61,480,832,3]` uint8; `action[60,9]` float32; 10 fps | `20260909-135841-262ee2`, all 5 cells passed | +| [Transfer](transfer/run_video_transfer_with_trt_llm.ipynb): edge | JSON/base64 `extra_params.edge`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | +| Transfer: blur | JSON/base64 `extra_params.blur`, `size=1104x832`, `format=mp4` | **1104×832**, 121 frames, 30 fps | `20260909-140456-faff27` | +| Transfer: depth | JSON/base64 `extra_params.depth`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | +| Transfer: segmentation | JSON/base64 `extra_params.seg`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | +| Transfer: WSM | JSON/base64 `extra_params.wsm`, `format=mp4` | 1280×720, 100 frames, 10 fps | `20260909-140456-faff27` | + +Both Action requests completed and decoded their tensors. The same run then +failed at the first Transfer request with HTTP 400 because `ffmpeg` was absent +from the server's `PATH`. After exposing the installed imageio-ffmpeg binary, +the full 12-cell Transfer notebook passed, with five HTTP 200 `video/mp4` +responses. Artifact validation run `20260909-150023-120a4f` decoded all seven +MP4s and checked the two safetensors payloads. The failed combined run is not +being counted as a successful Transfer run. + +The blur server log records `height=832 width=1104`, followed by +`Cosmos3 transfer target=1104x832 (WxH)` and HTTP 200. Its MP4 is 688,026 bytes, +SHA-256 `f5dc9767915a37f9b3306fe2a55b15984589b025d01698cd843be5ff2fa8b21d`. +The old 1280×720 summary describes an earlier request, not this notebook. + +For locating the exact retained evidence, gpu-run IDs refer to the author's +local `~/.config/gpu-run/runs//log` records. Videos, tensors, and Cosmos +Framework comparisons are under +`~/sim/cosmos3/transfer-action-output/matched-runs//`, with +`tensorrt-llm.mp4` and `cosmos-framework.mp4` in each case directory. +These local artifacts are not public downloads. + +The notebook file SHA-256 values at the tested cookbook commit are: + +| Notebook | SHA-256 | +| --- | --- | +| `action/run_fd_with_trt_llm.ipynb` | `13a243a6adfb1dba26c80c500aedf661ac7024edbf1823adfdde77f1eb17c885` | +| `action/run_id_with_trt_llm.ipynb` | `86232889e69e65dd73c1cc895896375ed586bbbc33ee15553511d2bfaaee2363` | +| `transfer/run_video_transfer_with_trt_llm.ipynb` | `61f960a34ae9e84e2eea7fa287df6d7cdae2206fa3b7c46a119f2e6d9651b63c` | + +## Runtime issues and remaining limits + +The [Mon 7 Sep 2026 merge hold](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5571009086) +was followed by the fixes and checks below. The evidence does **not** establish +that every issue is resolved in every release image. + +| Reported issue | Evidence and current scope | +| --- | --- | +| Forward-dynamics visual corruption on the reviewer's GB200 runtime | Later H100 and GB200 runs produced coherent output. The native GB200 run `20260910-180258-e3d281` executed all FD cells, used default GPU RNG, and produced 61 frames at 832×480/10 fps. It used the ARM64 rc26 wheel plus the upstream action image-bytes decoder fix, not the reviewer's exact daily container. The original corruption was not reproduced, and its root cause is not established. | +| Transfer videos did not play in Firefox | The notebook explicitly requests MP4, checks the response MIME type and MP4 signature, and embeds the result. All five later Transfer responses were MP4. Actual Firefox playback was not tested in these runs. | +| Distilled I2V entered the LLM/MPI launcher and failed to spawn | The documented launch explicitly selects VisualGen with `--visual_gen_args`. The tested Nano server launched one VisualGen worker and served requests successfully. This does not establish a fix for generic LLM/MPI spawning or certify distilled I2V. | +| Distilled T2I/I2V | Removed from this PR in [7ff4a86](https://github.com/NVIDIA/cosmos/commit/7ff4a8686f52e6fdd32ebf449171fec9b0695444). They are excluded, not reported as passing. | +| Guardrails | The requests set `use_guardrails=True`, but retained logs warn `No safety models found, returning safe`. These runs validate generation, not functioning safety checks. The [Thu 17 Sep 2026 follow-up](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5722101980) additionally reports HTTP 500 from `pathsec` with Guardrail cache symlinks and success only with guardrails disabled. Working guardrails remain unverified here. | + +The Thu 17 Sep follow-up separately reports successful FD generation on H100 +and H200 with rc26/rc27 images and a source overlay. It explicitly did not rerun +inverse dynamics, Transfer, or GB200; those results must not be extended to +those cases. + +This validation matrix covers **Nano** Action and Transfer. It does not certify +the shared Super launch, every audiovisual example, optional raw-source control +preprocessing, or later TensorRT-LLM revisions. The evidence supports the +documented request contract and the dimensions above, but not a blanket removal +of the merge hold while the guardrail/default-environment issue remains open. From e4d59b60ba8bd2cc86b4ed0796d3147aa7bfa01b Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 18 Sep 2026 08:23:49 -0700 Subject: [PATCH 12/14] docs: scope September merge hold to removed distilled models Signed-off-by: Igor Shovkun --- .../cosmos3/generator/trtllm-validation.md | 24 ++++++++++++------- 1 file changed, 16 insertions(+), 8 deletions(-) diff --git a/cookbooks/cosmos3/generator/trtllm-validation.md b/cookbooks/cosmos3/generator/trtllm-validation.md index 09524cb5..0ac0bb7c 100644 --- a/cookbooks/cosmos3/generator/trtllm-validation.md +++ b/cookbooks/cosmos3/generator/trtllm-validation.md @@ -84,18 +84,26 @@ The notebook file SHA-256 values at the tested cookbook commit are: | `action/run_id_with_trt_llm.ipynb` | `86232889e69e65dd73c1cc895896375ed586bbbc33ee15553511d2bfaaee2363` | | `transfer/run_video_transfer_with_trt_llm.ipynb` | `61f960a34ae9e84e2eea7fa287df6d7cdae2206fa3b7c46a119f2e6d9651b63c` | -## Runtime issues and remaining limits +## September 7 merge-hold issue -The [Mon 7 Sep 2026 merge hold](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5571009086) -was followed by the fixes and checks below. The evidence does **not** establish -that every issue is resolved in every release image. +The runtime issue behind the +[Mon 7 Sep 2026 merge hold](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5571009086) +concerned the distilled model. The distilled T2I/I2V examples were removed from +this PR in [7ff4a86](https://github.com/NVIDIA/cosmos/commit/7ff4a8686f52e6fdd32ebf449171fec9b0695444), +so that issue no longer applies to this PR's scope. This addresses the review +concern by excluding the affected examples; it does not claim an upstream fix +or a passing distilled-model test. + +## Separate runtime observations and validation limits + +The following observations qualify the retained generation evidence. They are +separate from the distilled-model issue behind the September 7 merge hold. | Reported issue | Evidence and current scope | | --- | --- | | Forward-dynamics visual corruption on the reviewer's GB200 runtime | Later H100 and GB200 runs produced coherent output. The native GB200 run `20260910-180258-e3d281` executed all FD cells, used default GPU RNG, and produced 61 frames at 832×480/10 fps. It used the ARM64 rc26 wheel plus the upstream action image-bytes decoder fix, not the reviewer's exact daily container. The original corruption was not reproduced, and its root cause is not established. | | Transfer videos did not play in Firefox | The notebook explicitly requests MP4, checks the response MIME type and MP4 signature, and embeds the result. All five later Transfer responses were MP4. Actual Firefox playback was not tested in these runs. | -| Distilled I2V entered the LLM/MPI launcher and failed to spawn | The documented launch explicitly selects VisualGen with `--visual_gen_args`. The tested Nano server launched one VisualGen worker and served requests successfully. This does not establish a fix for generic LLM/MPI spawning or certify distilled I2V. | -| Distilled T2I/I2V | Removed from this PR in [7ff4a86](https://github.com/NVIDIA/cosmos/commit/7ff4a8686f52e6fdd32ebf449171fec9b0695444). They are excluded, not reported as passing. | +| VisualGen launch | The documented launch explicitly selects VisualGen with `--visual_gen_args`. The tested Nano server launched one VisualGen worker and served requests successfully. | | Guardrails | The requests set `use_guardrails=True`, but retained logs warn `No safety models found, returning safe`. These runs validate generation, not functioning safety checks. The [Thu 17 Sep 2026 follow-up](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5722101980) additionally reports HTTP 500 from `pathsec` with Guardrail cache symlinks and success only with guardrails disabled. Working guardrails remain unverified here. | The Thu 17 Sep follow-up separately reports successful FD generation on H100 @@ -106,5 +114,5 @@ those cases. This validation matrix covers **Nano** Action and Transfer. It does not certify the shared Super launch, every audiovisual example, optional raw-source control preprocessing, or later TensorRT-LLM revisions. The evidence supports the -documented request contract and the dimensions above, but not a blanket removal -of the merge hold while the guardrail/default-environment issue remains open. +documented generation request contract and the dimensions above; it does not +certify functioning guardrails. From 1595ad3e2cbaedb3bc7111de273cee4a846f552c Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Fri, 18 Sep 2026 21:54:07 -0700 Subject: [PATCH 13/14] docs(cosmos3): validate server-side NLTK guardrail workaround --- cookbooks/cosmos3/README.md | 46 +++++++++ cookbooks/cosmos3/generator/action/README.md | 4 +- .../action/run_fd_with_trt_llm.ipynb | 1 + .../action/run_id_with_trt_llm.ipynb | 1 + .../cosmos3/generator/transfer/README.md | 4 +- .../run_video_transfer_with_trt_llm.ipynb | 3 +- .../cosmos3/generator/trtllm-validation.md | 93 +++++++++++++++++-- 7 files changed, 143 insertions(+), 9 deletions(-) diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 4a27575b..e92cf396 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -238,6 +238,52 @@ pip uninstall -y opencv-python pip install opencv-python-headless==5.0.0.93 ``` +#### Server-side NLTK data setup + +With `cosmos_guardrail==0.3.0`, NLTK's path checks can reject tokenizer or +dictionary files that are symlinks from a Hugging Face snapshot into its +`blobs/` directory. Prepare a separate copy of the NLTK data as regular files +**on the server, in the same container and shell used to launch TensorRT-LLM**: + +```bash +export NLTK_DATA="$(mktemp -d "${TMPDIR:-/tmp}/cosmos3-nltk.XXXXXX")" +python3 - <<'PY' +import os +import shutil +from pathlib import Path + +import nltk +from huggingface_hub import snapshot_download + +snapshot = snapshot_download( + "nvidia/Cosmos-1.0-Guardrail", + revision="cf03c0395fac8c4de386c0bdab12cc4fc8d66362", + allow_patterns=["blocklist/**"], +) +source = Path(snapshot) / "blocklist" / "nltk_data" +destination = Path(os.environ["NLTK_DATA"]).resolve() +shutil.copytree(source, destination, symlinks=False, dirs_exist_ok=True) +assert not any(path.is_symlink() for path in destination.rglob("*")) + +# Check both resource lookups used by the text blocklist before starting a GPU server. +nltk.data.path[:] = [str(destination)] +tokens = nltk.word_tokenize("You are an autonomous vehicle planning system.") +assert nltk.WordNetLemmatizer().lemmatize("vehicles") == "vehicle" +print("Guardrail NLTK data ready:", destination, tokens) +PY +``` + +Keep `NLTK_DATA` exported when starting the server below. Repeat this setup +after recreating the container or removing the temporary directory. Running it +only in the client notebook's environment does not configure a remote server. +This workaround does not rewrite cached files or symlinks, and keeps NLTK +path security and `use_guardrails=True` enabled. + +The separate `No safety models found, returning safe` warning in guardrail +0.3.0 refers to its intentionally empty video-content classifier list. Text +checks and face blurring remain configured; the NLTK workaround does not enable +video-content classification or suppress that warning. + Set the TensorRT-LLM source root for the shared VisualGen config YAMLs: ```bash diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 63e160f3..b93e688d 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -174,7 +174,9 @@ have available. ### Quickstart Set up and launch the VisualGen server with the one-GPU Nano config: -[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). Action requests +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), including the +[server-side NLTK data setup](../../README.md#server-side-nltk-data-setup). +Keep `NLTK_DATA` exported in the server's shell. Action requests upload a conditioning image as multipart `image_reference` or a video as `video_reference`, and put the mode-specific fields under `extra_params`. Use the same trained AV prompt as the Cosmos Framework reference; the legacy diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb index db73988d..0be2bbf9 100644 --- a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -30,6 +30,7 @@ "\n", "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", "\n", "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb index 3fa07636..347aa9a7 100644 --- a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -29,6 +29,7 @@ "\n", "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", "\n", "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index ed7936c3..690dbaee 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -199,7 +199,9 @@ reuses the previews from [`preview_helpers.py`](./preview_helpers.py), writing o ### Quickstart Set up and launch Nano or Super as described in the shared -[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator). TensorRT-LLM +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), including the +[server-side NLTK data setup](../../README.md#server-side-nltk-data-setup). +Keep `NLTK_DATA` exported in the server's shell. TensorRT-LLM accepts a raw source video as multipart `video_reference` when edge or blur will be computed on the server. The checked-in assets are already precomputed controls, so this example base64-encodes the control inside its hint and does diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb index 7183a6ef..a5955f17 100644 --- a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -29,7 +29,8 @@ "## Start the TensorRT-LLM server\n", "\n", "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then launch one\n", - "server from the TensorRT-LLM checkout:\n", + "server from the TensorRT-LLM checkout.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", "\n", "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", diff --git a/cookbooks/cosmos3/generator/trtllm-validation.md b/cookbooks/cosmos3/generator/trtllm-validation.md index 0ac0bb7c..1e6dfb8b 100644 --- a/cookbooks/cosmos3/generator/trtllm-validation.md +++ b/cookbooks/cosmos3/generator/trtllm-validation.md @@ -2,12 +2,13 @@ This record reconciles the saved GPU runs with the notebook requests in [cookbook commit `7f973ea81f4b37bab2d875fa34a90313c35a9990`](https://github.com/NVIDIA/cosmos/commit/7f973ea81f4b37bab2d875fa34a90313c35a9990). -It was audited on Fri 18 Sep 2026 (Pacific) using the retained logs and -artifacts; this documentation update did not run new GPU inference. +The historical runs were audited on Fri 18 Sep 2026 (Pacific) using retained +logs and artifacts. A fresh run of the NLTK workaround is recorded separately +below; do not confuse its environment with the earlier runs. ## Compatible runtime -The Action and Nano Transfer notebook runs used TensorRT-LLM source revision +The historical Action and Nano Transfer notebook runs used TensorRT-LLM source revision [`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e) in a native Python 3.12 environment, with PyTorch `2.12.0+cu130` and `cosmos_guardrail==0.3.0`. The runtime reported `1.3.0rc26`; that package @@ -84,6 +85,84 @@ The notebook file SHA-256 values at the tested cookbook commit are: | `action/run_id_with_trt_llm.ipynb` | `86232889e69e65dd73c1cc895896375ed586bbbc33ee15553511d2bfaaee2363` | | `transfer/run_video_transfer_with_trt_llm.ipynb` | `61f960a34ae9e84e2eea7fa287df6d7cdae2206fa3b7c46a119f2e6d9651b63c` | +## NLTK workaround validation — Fri 18 Sep 2026 + +The [server-side NLTK setup](../README.md#server-side-nltk-data-setup) copies +the Guardrail snapshot's NLTK resources into a private directory as regular +files, then exports `NLTK_DATA` in the server shell. It does not disable NLTK +path security, change request guardrail settings, or patch either library. + +The runtime was rebuilt natively from the same TensorRT-LLM source revision +`bca6761ab84fbcd58fc7f914eade7de48b32e35e`, using one H200 on `ipp2-0039`. +Build run `20260918-203438-1aaf41` passed and installed +`1.3.0rc26+bca6761ab8`. This is a native build, not a container validation. +The environment used Python 3.12.3, PyTorch `2.12.0+cu130`, Transformers +`5.5.4`, `cosmos_guardrail==0.3.0`, NLTK `3.10.3`, and driver `595.58.03`. +The installed imageio-ffmpeg 7.0.2 executable was exposed as `ffmpeg` in the +runtime's own `PATH`; no system package installation was performed. + +Preflight run `20260918-205406-56f855` executed the README's NLTK setup +verbatim, then called the installed `Blocklist.is_safe()` on the FD prompt. +It passed with `nltk.pathsec.ENFORCE=True`, 383 blocklist entries and 1,430 +exact-match entries loaded, and the tokenizer resolved inside the new NLTK +directory. This preflight is a CPU component check, not GPU inference. + +Notebook run `20260918-205707-474ba5` executed the README setup and the FD +notebook's server command verbatim, then executed the checked-in code cells +of both Action notebooks and the full Transfer notebook. All requests retain +`use_guardrails=True`. It **passed**, exiting 0 after 3,149.8 seconds +(20:57–21:49 Pacific). All 5 FD, 5 ID, and 12 Transfer code cells executed +without cell errors; all seven requests returned HTTP 200. There was no +`pathsec` exception or HTTP 500. Notebook inference code was not changed for +this workaround; only setup instructions in Markdown cells were added. + +Artifact-check run `20260918-215026-556c5a` decoded every frame of every MP4 +and verified the following results. Both Action safetensors files contain +`video[61,480,832,3]` uint8 and finite `action[60,9]` float32 tensors. + +| Case | Result | Decoded MP4 (width × height, frames, fps) | +| --- | --- | --- | +| Forward dynamics | PASS | 832×480, 61, 10 | +| Inverse dynamics | PASS | 832×480, 61, 10 | +| Transfer edge | PASS | 1280×720, 121, 30 | +| Transfer blur | PASS | 1104×832, 121, 30 | +| Transfer depth | PASS | 1280×720, 121, 30 | +| Transfer segmentation | PASS | 1280×720, 121, 30 | +| Transfer WSM | PASS | 1280×720, 100, 10 | + +The executed notebook SHA-256 values identify the tested files independently +of later edits to this validation record: + +| Notebook | SHA-256 | +| --- | --- | +| `action/run_fd_with_trt_llm.ipynb` | `2ee065cf6a96fdfdd64c00f2f89d972e41d2c198d1e03f6f0dbe8e304780cf52` | +| `action/run_id_with_trt_llm.ipynb` | `1f7d5c0e46ca19350cc7fb3a001ac774308a5f5762dcab55896065639a1f2638` | +| `transfer/run_video_transfer_with_trt_llm.ipynb` | `8946c3fdd1fb81bdde53bff67e76527cf693bf49155db832901a4fb4d3ec24d3` | + +Full logs, dependency versions, executed notebooks, tensors, and artifact +hashes are retained under the author's scratch directory +`/home/scratch.ishovkun_gpu/gpu-run/artifacts/pr310-guardrail-20260918-2057/`. +Local videos are under +`~/sim/cosmos3/transfer-action-output/guardrail-workaround-20260918//`, +named `tensorrt-llm.mp4`, beside the previously saved `cosmos-framework.mp4` +reference. These are local evidence paths, not public downloads. The Framework +references were not rerun; visual comparison is not a pixel-parity claim. + +The documented `nvidia/Cosmos3-Nano` model ID resolved to snapshot +`e59a53c25979a090fa8706c9acc0c254a6e89b92`. A read-only cache comparison +(`20260918-211533-9a4325`) confirmed that its model safetensors files share +the same cached data as the earlier `7a312c...` snapshot. The snapshot IDs +are nevertheless recorded separately. + +This is not a warning-free runtime. Guardrail 0.3.0 still warns about its +intentionally disabled video-content classifier; its text checks and +RetinaFace postprocessing remain configured. Native MPI optional-plugin, +torchao compatibility, and upstream deprecation/shape-warmup warnings are +retained in the logs. The notebook execution harness also reported an +IPython history-file error on shared storage, and Python reported a leaked +semaphore at shutdown. None is silently filtered out or counted as validation +of the disabled video-content classifier. + ## September 7 merge-hold issue The runtime issue behind the @@ -104,7 +183,7 @@ separate from the distilled-model issue behind the September 7 merge hold. | Forward-dynamics visual corruption on the reviewer's GB200 runtime | Later H100 and GB200 runs produced coherent output. The native GB200 run `20260910-180258-e3d281` executed all FD cells, used default GPU RNG, and produced 61 frames at 832×480/10 fps. It used the ARM64 rc26 wheel plus the upstream action image-bytes decoder fix, not the reviewer's exact daily container. The original corruption was not reproduced, and its root cause is not established. | | Transfer videos did not play in Firefox | The notebook explicitly requests MP4, checks the response MIME type and MP4 signature, and embeds the result. All five later Transfer responses were MP4. Actual Firefox playback was not tested in these runs. | | VisualGen launch | The documented launch explicitly selects VisualGen with `--visual_gen_args`. The tested Nano server launched one VisualGen worker and served requests successfully. | -| Guardrails | The requests set `use_guardrails=True`, but retained logs warn `No safety models found, returning safe`. These runs validate generation, not functioning safety checks. The [Thu 17 Sep 2026 follow-up](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5722101980) additionally reports HTTP 500 from `pathsec` with Guardrail cache symlinks and success only with guardrails disabled. Working guardrails remain unverified here. | +| Guardrails | The requests set `use_guardrails=True`. In `cosmos_guardrail==0.3.0`, `No safety models found, returning safe` comes from the intentionally empty video-content classifier list, not from failed loading of every safety component. Its constructor configures Blocklist and Qwen3Guard text checks plus RetinaFace face blurring. The separate [Thu 17 Sep 2026 follow-up](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5722101980) reports HTTP 500 from `pathsec` with Guardrail cache symlinks. The [server-side NLTK setup](../README.md#server-side-nltk-data-setup) materializes those resources without disabling path security. The fresh run above passed both Action notebooks and all five Transfer cases with that workaround and the unmodified guardrail 0.3.0 package. The disabled video-content classifier remains unvalidated. | The Thu 17 Sep follow-up separately reports successful FD generation on H100 and H200 with rc26/rc27 images and a source overlay. It explicitly did not rerun @@ -114,5 +193,7 @@ those cases. This validation matrix covers **Nano** Action and Transfer. It does not certify the shared Super launch, every audiovisual example, optional raw-source control preprocessing, or later TensorRT-LLM revisions. The evidence supports the -documented generation request contract and the dimensions above; it does not -certify functioning guardrails. +documented generation request contract, the dimensions above, and a working +default guardrail 0.3.0 configuration with the NLTK workaround. It is not a +safety-efficacy benchmark and does not certify the disabled video-content +classifier. From b7092b5c8cfd48b52d58bd0236b66114615284e0 Mon Sep 17 00:00:00 2001 From: Igor Shovkun Date: Mon, 28 Sep 2026 10:54:33 -0700 Subject: [PATCH 14/14] Remove historical TensorRT-LLM validation record from cookbook --- cookbooks/cosmos3/README.md | 5 +- cookbooks/cosmos3/generator/action/README.md | 4 +- .../cosmos3/generator/transfer/README.md | 6 - .../cosmos3/generator/trtllm-validation.md | 199 ------------------ 4 files changed, 2 insertions(+), 212 deletions(-) delete mode 100644 cookbooks/cosmos3/generator/trtllm-validation.md diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 1e8eeb1c..9dd21aa5 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -185,10 +185,7 @@ These changes are merged on TensorRT-LLM `main`. The Action and Transfer notebooks were executed against source revision [`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e). Use that revision to reproduce their request contract, or a newer build with -the same API. The older `799d7d42` validation used `input_reference` and does -not validate these notebooks' `image_reference` / `video_reference` fields. -See the [Action and Transfer validation record](generator/trtllm-validation.md) -for the executed cases, output dimensions, and remaining runtime limitations. +the same API. The source revision is significant: a package version of `1.3.0rc26` alone does not establish compatibility with the action image-decoding path. diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index b93e688d..8cd94264 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -184,9 +184,7 @@ Use the same trained AV prompt as the Cosmos Framework reference; the legacy Both notebooks' current multipart fields were exercised against TensorRT-LLM source revision `bca6761ab84fbcd58fc7f914eade7de48b32e35e`. Follow the source -pin in the shared setup; the earlier `799d7d42` run used the old -`input_reference` API. The [validation record](../trtllm-validation.md) covers -forward and inverse dynamics separately and records the guardrail limitation. +pin in the shared setup. ```python import json diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index 690dbaee..735c776a 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -276,12 +276,6 @@ requests its native 4:3 bucket, 1104×832; the other examples request at 10 fps. `/v1/videos/generations` remains only as a deprecated alias of the canonical blocking `/v1/videos/sync` route. -The [validation record](../trtllm-validation.md) reports the later execution of -all five controls at TensorRT-LLM source revision -`bca6761ab84fbcd58fc7f914eade7de48b32e35e`, including blur at 1104×832. -The earlier 1280×720 blur result used a previous request and does not describe -the checked-in notebook. - ### TensorRT-LLM notebook walkthrough [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) diff --git a/cookbooks/cosmos3/generator/trtllm-validation.md b/cookbooks/cosmos3/generator/trtllm-validation.md deleted file mode 100644 index 1e6dfb8b..00000000 --- a/cookbooks/cosmos3/generator/trtllm-validation.md +++ /dev/null @@ -1,199 +0,0 @@ -# TensorRT-LLM Action and Transfer validation - -This record reconciles the saved GPU runs with the notebook requests in -[cookbook commit `7f973ea81f4b37bab2d875fa34a90313c35a9990`](https://github.com/NVIDIA/cosmos/commit/7f973ea81f4b37bab2d875fa34a90313c35a9990). -The historical runs were audited on Fri 18 Sep 2026 (Pacific) using retained -logs and artifacts. A fresh run of the NLTK workaround is recorded separately -below; do not confuse its environment with the earlier runs. - -## Compatible runtime - -The historical Action and Nano Transfer notebook runs used TensorRT-LLM source revision -[`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e) -in a native Python 3.12 environment, with PyTorch `2.12.0+cu130` and -`cosmos_guardrail==0.3.0`. The runtime reported `1.3.0rc26`; that package -version alone does not identify the source revision. These were native runs, -so there is no container digest. The retained evidence does not establish a -fresh C++ rebuild or an immutable fingerprint of every installed dependency. - -At this revision, the -[request schema](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/openai_protocol.py#L2067) -and [multipart parser](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/openai_video_routes.py#L311) -accept `image_reference` and `video_reference`. -The [conversion layer](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/serve/visual_gen_utils.py#L482) -passes uploaded media to the worker, and the -[Action image decoder](https://github.com/NVIDIA/TensorRT-LLM/blob/bca6761ab84fbcd58fc7f914eade7de48b32e35e/tensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.py#L364) -handles encoded image bytes. The earlier `799d7d42` runs used `input_reference`; -they are historical evidence, not validation of the current request format. - -The server used the checked-in `cosmos3-nano-1gpu.yaml`: - -```bash -trtllm-serve nvidia/Cosmos3-Nano \ - --visual_gen_args examples/visual_gen/configs/cosmos3-nano-1gpu.yaml \ - --host 127.0.0.1 --port 8000 -``` - -The checkpoint resolved to snapshot -`7a312c868bcce8e40b3eb40861300a9d0ba3fde1`. Each run used one H100 80GB on -`ipp2-0160`. See the [shared setup](../README.md#tensorrt-llm-generator) for -the source pin and launch commands. - -## Executed notebook requests - -The runs below executed the checked-in code cells, including request creation -and response decoding. Action used the checked-in AV assets, prompt -`You are an autonomous vehicle planning system.`, and seed 0. Transfer used -the checked-in controls and captions, seed 2026, 50 steps, and flow shift 10. -Dimensions in this table are width × height. - -| Notebook/case | Request to `POST /v1/videos/sync` | Decoded result | GPU evidence | -| --- | --- | --- | --- | -| [Forward dynamics](action/run_fd_with_trt_llm.ipynb) | Multipart `image_reference`, `format=safetensors` | `video[61,480,832,3]` uint8; `action[60,9]` float32; 10 fps | `20260909-135841-262ee2`, all 5 cells passed | -| [Inverse dynamics](action/run_id_with_trt_llm.ipynb) | Multipart `video_reference`, `format=safetensors` | `video[61,480,832,3]` uint8; `action[60,9]` float32; 10 fps | `20260909-135841-262ee2`, all 5 cells passed | -| [Transfer](transfer/run_video_transfer_with_trt_llm.ipynb): edge | JSON/base64 `extra_params.edge`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | -| Transfer: blur | JSON/base64 `extra_params.blur`, `size=1104x832`, `format=mp4` | **1104×832**, 121 frames, 30 fps | `20260909-140456-faff27` | -| Transfer: depth | JSON/base64 `extra_params.depth`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | -| Transfer: segmentation | JSON/base64 `extra_params.seg`, `format=mp4` | 1280×720, 121 frames, 30 fps | `20260909-140456-faff27` | -| Transfer: WSM | JSON/base64 `extra_params.wsm`, `format=mp4` | 1280×720, 100 frames, 10 fps | `20260909-140456-faff27` | - -Both Action requests completed and decoded their tensors. The same run then -failed at the first Transfer request with HTTP 400 because `ffmpeg` was absent -from the server's `PATH`. After exposing the installed imageio-ffmpeg binary, -the full 12-cell Transfer notebook passed, with five HTTP 200 `video/mp4` -responses. Artifact validation run `20260909-150023-120a4f` decoded all seven -MP4s and checked the two safetensors payloads. The failed combined run is not -being counted as a successful Transfer run. - -The blur server log records `height=832 width=1104`, followed by -`Cosmos3 transfer target=1104x832 (WxH)` and HTTP 200. Its MP4 is 688,026 bytes, -SHA-256 `f5dc9767915a37f9b3306fe2a55b15984589b025d01698cd843be5ff2fa8b21d`. -The old 1280×720 summary describes an earlier request, not this notebook. - -For locating the exact retained evidence, gpu-run IDs refer to the author's -local `~/.config/gpu-run/runs//log` records. Videos, tensors, and Cosmos -Framework comparisons are under -`~/sim/cosmos3/transfer-action-output/matched-runs//`, with -`tensorrt-llm.mp4` and `cosmos-framework.mp4` in each case directory. -These local artifacts are not public downloads. - -The notebook file SHA-256 values at the tested cookbook commit are: - -| Notebook | SHA-256 | -| --- | --- | -| `action/run_fd_with_trt_llm.ipynb` | `13a243a6adfb1dba26c80c500aedf661ac7024edbf1823adfdde77f1eb17c885` | -| `action/run_id_with_trt_llm.ipynb` | `86232889e69e65dd73c1cc895896375ed586bbbc33ee15553511d2bfaaee2363` | -| `transfer/run_video_transfer_with_trt_llm.ipynb` | `61f960a34ae9e84e2eea7fa287df6d7cdae2206fa3b7c46a119f2e6d9651b63c` | - -## NLTK workaround validation — Fri 18 Sep 2026 - -The [server-side NLTK setup](../README.md#server-side-nltk-data-setup) copies -the Guardrail snapshot's NLTK resources into a private directory as regular -files, then exports `NLTK_DATA` in the server shell. It does not disable NLTK -path security, change request guardrail settings, or patch either library. - -The runtime was rebuilt natively from the same TensorRT-LLM source revision -`bca6761ab84fbcd58fc7f914eade7de48b32e35e`, using one H200 on `ipp2-0039`. -Build run `20260918-203438-1aaf41` passed and installed -`1.3.0rc26+bca6761ab8`. This is a native build, not a container validation. -The environment used Python 3.12.3, PyTorch `2.12.0+cu130`, Transformers -`5.5.4`, `cosmos_guardrail==0.3.0`, NLTK `3.10.3`, and driver `595.58.03`. -The installed imageio-ffmpeg 7.0.2 executable was exposed as `ffmpeg` in the -runtime's own `PATH`; no system package installation was performed. - -Preflight run `20260918-205406-56f855` executed the README's NLTK setup -verbatim, then called the installed `Blocklist.is_safe()` on the FD prompt. -It passed with `nltk.pathsec.ENFORCE=True`, 383 blocklist entries and 1,430 -exact-match entries loaded, and the tokenizer resolved inside the new NLTK -directory. This preflight is a CPU component check, not GPU inference. - -Notebook run `20260918-205707-474ba5` executed the README setup and the FD -notebook's server command verbatim, then executed the checked-in code cells -of both Action notebooks and the full Transfer notebook. All requests retain -`use_guardrails=True`. It **passed**, exiting 0 after 3,149.8 seconds -(20:57–21:49 Pacific). All 5 FD, 5 ID, and 12 Transfer code cells executed -without cell errors; all seven requests returned HTTP 200. There was no -`pathsec` exception or HTTP 500. Notebook inference code was not changed for -this workaround; only setup instructions in Markdown cells were added. - -Artifact-check run `20260918-215026-556c5a` decoded every frame of every MP4 -and verified the following results. Both Action safetensors files contain -`video[61,480,832,3]` uint8 and finite `action[60,9]` float32 tensors. - -| Case | Result | Decoded MP4 (width × height, frames, fps) | -| --- | --- | --- | -| Forward dynamics | PASS | 832×480, 61, 10 | -| Inverse dynamics | PASS | 832×480, 61, 10 | -| Transfer edge | PASS | 1280×720, 121, 30 | -| Transfer blur | PASS | 1104×832, 121, 30 | -| Transfer depth | PASS | 1280×720, 121, 30 | -| Transfer segmentation | PASS | 1280×720, 121, 30 | -| Transfer WSM | PASS | 1280×720, 100, 10 | - -The executed notebook SHA-256 values identify the tested files independently -of later edits to this validation record: - -| Notebook | SHA-256 | -| --- | --- | -| `action/run_fd_with_trt_llm.ipynb` | `2ee065cf6a96fdfdd64c00f2f89d972e41d2c198d1e03f6f0dbe8e304780cf52` | -| `action/run_id_with_trt_llm.ipynb` | `1f7d5c0e46ca19350cc7fb3a001ac774308a5f5762dcab55896065639a1f2638` | -| `transfer/run_video_transfer_with_trt_llm.ipynb` | `8946c3fdd1fb81bdde53bff67e76527cf693bf49155db832901a4fb4d3ec24d3` | - -Full logs, dependency versions, executed notebooks, tensors, and artifact -hashes are retained under the author's scratch directory -`/home/scratch.ishovkun_gpu/gpu-run/artifacts/pr310-guardrail-20260918-2057/`. -Local videos are under -`~/sim/cosmos3/transfer-action-output/guardrail-workaround-20260918//`, -named `tensorrt-llm.mp4`, beside the previously saved `cosmos-framework.mp4` -reference. These are local evidence paths, not public downloads. The Framework -references were not rerun; visual comparison is not a pixel-parity claim. - -The documented `nvidia/Cosmos3-Nano` model ID resolved to snapshot -`e59a53c25979a090fa8706c9acc0c254a6e89b92`. A read-only cache comparison -(`20260918-211533-9a4325`) confirmed that its model safetensors files share -the same cached data as the earlier `7a312c...` snapshot. The snapshot IDs -are nevertheless recorded separately. - -This is not a warning-free runtime. Guardrail 0.3.0 still warns about its -intentionally disabled video-content classifier; its text checks and -RetinaFace postprocessing remain configured. Native MPI optional-plugin, -torchao compatibility, and upstream deprecation/shape-warmup warnings are -retained in the logs. The notebook execution harness also reported an -IPython history-file error on shared storage, and Python reported a leaked -semaphore at shutdown. None is silently filtered out or counted as validation -of the disabled video-content classifier. - -## September 7 merge-hold issue - -The runtime issue behind the -[Mon 7 Sep 2026 merge hold](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5571009086) -concerned the distilled model. The distilled T2I/I2V examples were removed from -this PR in [7ff4a86](https://github.com/NVIDIA/cosmos/commit/7ff4a8686f52e6fdd32ebf449171fec9b0695444), -so that issue no longer applies to this PR's scope. This addresses the review -concern by excluding the affected examples; it does not claim an upstream fix -or a passing distilled-model test. - -## Separate runtime observations and validation limits - -The following observations qualify the retained generation evidence. They are -separate from the distilled-model issue behind the September 7 merge hold. - -| Reported issue | Evidence and current scope | -| --- | --- | -| Forward-dynamics visual corruption on the reviewer's GB200 runtime | Later H100 and GB200 runs produced coherent output. The native GB200 run `20260910-180258-e3d281` executed all FD cells, used default GPU RNG, and produced 61 frames at 832×480/10 fps. It used the ARM64 rc26 wheel plus the upstream action image-bytes decoder fix, not the reviewer's exact daily container. The original corruption was not reproduced, and its root cause is not established. | -| Transfer videos did not play in Firefox | The notebook explicitly requests MP4, checks the response MIME type and MP4 signature, and embeds the result. All five later Transfer responses were MP4. Actual Firefox playback was not tested in these runs. | -| VisualGen launch | The documented launch explicitly selects VisualGen with `--visual_gen_args`. The tested Nano server launched one VisualGen worker and served requests successfully. | -| Guardrails | The requests set `use_guardrails=True`. In `cosmos_guardrail==0.3.0`, `No safety models found, returning safe` comes from the intentionally empty video-content classifier list, not from failed loading of every safety component. Its constructor configures Blocklist and Qwen3Guard text checks plus RetinaFace face blurring. The separate [Thu 17 Sep 2026 follow-up](https://github.com/NVIDIA/cosmos/pull/310#issuecomment-5722101980) reports HTTP 500 from `pathsec` with Guardrail cache symlinks. The [server-side NLTK setup](../README.md#server-side-nltk-data-setup) materializes those resources without disabling path security. The fresh run above passed both Action notebooks and all five Transfer cases with that workaround and the unmodified guardrail 0.3.0 package. The disabled video-content classifier remains unvalidated. | - -The Thu 17 Sep follow-up separately reports successful FD generation on H100 -and H200 with rc26/rc27 images and a source overlay. It explicitly did not rerun -inverse dynamics, Transfer, or GB200; those results must not be extended to -those cases. - -This validation matrix covers **Nano** Action and Transfer. It does not certify -the shared Super launch, every audiovisual example, optional raw-source control -preprocessing, or later TensorRT-LLM revisions. The evidence supports the -documented generation request contract, the dimensions above, and a working -default guardrail 0.3.0 configuration with the NLTK workaround. It is not a -safety-efficacy benchmark and does not certify the disabled video-content -classifier.