diff --git a/cookbooks/cosmos3/README.md b/cookbooks/cosmos3/README.md index 3565f13e..9dd21aa5 100644 --- a/cookbooks/cosmos3/README.md +++ b/cookbooks/cosmos3/README.md @@ -8,7 +8,7 @@ backend you want to run and follow that one section. | --- | --- | --- | | [Cosmos Framework](#cosmos-framework) | Native PyTorch inference, launched with `torchrun` | Reasoner, Generator (Audiovisual, Action, **Transfer**) | | [Diffusers](#diffusers) | Direct generation with `Cosmos3OmniPipeline` | Generator (Audiovisual) | -| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video/audio generation) | Generator (Audiovisual) | +| [TensorRT-LLM Generator](#tensorrt-llm-generator) | OpenAI-compatible VisualGen server (image/video/audio/action/transfer generation) | Generator (Audiovisual, Action, **Transfer**) | | [TensorRT-LLM Reasoner](#tensorrt-llm-reasoner) | OpenAI-compatible image/video reasoning server | Reasoner | | [Transformers](#transformers) | Hugging Face Transformers inference | Reasoner | | [vLLM](#vllm) | OpenAI-compatible reasoning server (image/video understanding) | Reasoner | @@ -168,17 +168,26 @@ uv pip install --torch-backend=cu130 \ ## TensorRT-LLM Generator OpenAI-compatible **VisualGen** server for Generator audiovisual text-to-image, -text-to-video, image-to-video, video-to-video, and synchronized audio examples. +text-to-video, image-to-video, video-to-video, synchronized audio, Transfer, and +Action examples. Initial Cosmos3 support was added in TensorRT-LLM PR [#14824](https://github.com/NVIDIA/TensorRT-LLM/pull/14824), synchronized audio in [#14827](https://github.com/NVIDIA/TensorRT-LLM/pull/14827), and -video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155). The -DMD2-distilled four-step checkpoints were added in +video-to-video in [#16155](https://github.com/NVIDIA/TensorRT-LLM/pull/16155), +Transfer in [#16394](https://github.com/NVIDIA/TensorRT-LLM/pull/16394), and +Action in [#17325](https://github.com/NVIDIA/TensorRT-LLM/pull/17325). +The DMD2-distilled four-step checkpoints were added in [#16563](https://github.com/NVIDIA/TensorRT-LLM/pull/16563) (text-to-image) and [#16690](https://github.com/NVIDIA/TensorRT-LLM/pull/16690) (image-to-video), and Cosmos3-Edge (Nemotron-dense backbone) in [#16773](https://github.com/NVIDIA/TensorRT-LLM/pull/16773). -Use a TensorRT-LLM checkout or package that includes those changes. +These changes are merged on TensorRT-LLM `main`. The Action and Transfer +notebooks were executed against source revision +[`bca6761ab84fbcd58fc7f914eade7de48b32e35e`](https://github.com/NVIDIA/TensorRT-LLM/commit/bca6761ab84fbcd58fc7f914eade7de48b32e35e). +Use that revision to reproduce their request contract, or a newer build with +the same API. +The source revision is significant: a package version of `1.3.0rc26` alone +does not establish compatibility with the action image-decoding path. Install TensorRT-LLM following its upstream documentation. @@ -189,11 +198,13 @@ Cosmos3 VisualGen change before it is available in your installed package or release image. ```bash -apt-get update && apt-get -y install ffmpeg git git-lfs +apt-get update && apt-get -y install git git-lfs git lfs install git clone https://github.com/NVIDIA/TensorRT-LLM.git cd TensorRT-LLM +# Source revision used for the Action/Transfer notebook validation below. +git checkout bca6761ab84fbcd58fc7f914eade7de48b32e35e git submodule update --init --recursive git lfs pull @@ -208,6 +219,7 @@ docker run --rm -it \ nvcr.io/nvidia/tensorrt-llm/devel: # Inside the container: +apt-get update && apt-get -y install ffmpeg python3 scripts/build_wheel.py --use_ccache --skip_building_wheel --linking_install_binary pip install -e . ``` @@ -221,10 +233,58 @@ explicitly disable guardrails before starting the server: ```bash pip install cosmos_guardrail==0.3.0 -# If needed by your OpenCV stack: -# pip uninstall opencv-python +# On headless servers without libGL.so.1, replace the OpenCV wheel pulled in by +# cosmos_guardrail with the matching headless build: +pip uninstall -y opencv-python +pip install opencv-python-headless==5.0.0.93 +``` + +#### Server-side NLTK data setup + +With `cosmos_guardrail==0.3.0`, NLTK's path checks can reject tokenizer or +dictionary files that are symlinks from a Hugging Face snapshot into its +`blobs/` directory. Prepare a separate copy of the NLTK data as regular files +**on the server, in the same container and shell used to launch TensorRT-LLM**: + +```bash +export NLTK_DATA="$(mktemp -d "${TMPDIR:-/tmp}/cosmos3-nltk.XXXXXX")" +python3 - <<'PY' +import os +import shutil +from pathlib import Path + +import nltk +from huggingface_hub import snapshot_download + +snapshot = snapshot_download( + "nvidia/Cosmos-1.0-Guardrail", + revision="cf03c0395fac8c4de386c0bdab12cc4fc8d66362", + allow_patterns=["blocklist/**"], +) +source = Path(snapshot) / "blocklist" / "nltk_data" +destination = Path(os.environ["NLTK_DATA"]).resolve() +shutil.copytree(source, destination, symlinks=False, dirs_exist_ok=True) +assert not any(path.is_symlink() for path in destination.rglob("*")) + +# Check both resource lookups used by the text blocklist before starting a GPU server. +nltk.data.path[:] = [str(destination)] +tokens = nltk.word_tokenize("You are an autonomous vehicle planning system.") +assert nltk.WordNetLemmatizer().lemmatize("vehicles") == "vehicle" +print("Guardrail NLTK data ready:", destination, tokens) +PY ``` +Keep `NLTK_DATA` exported when starting the server below. Repeat this setup +after recreating the container or removing the temporary directory. Running it +only in the client notebook's environment does not configure a remote server. +This workaround does not rewrite cached files or symlinks, and keeps NLTK +path security and `use_guardrails=True` enabled. + +The separate `No safety models found, returning safe` warning in guardrail +0.3.0 refers to its intentionally empty video-content classifier list. Text +checks and face blurring remain configured; the NLTK workaround does not enable +video-content classification or suppress that warning. + Set the TensorRT-LLM source root for the shared VisualGen config YAMLs. Run this from inside the TensorRT-LLM checkout — the directory the `git clone` above created, which is where `examples/` lives — or point `TRTLLM_ROOT` at that @@ -304,22 +364,42 @@ runs both against a running server. These students cover text-to-image and image-to-video only; use the base checkpoints for text-to-video, video-to-video, and synchronized audio. -The server exposes `/health`, `/v1/videos/generations`, `/v1/videos`, and -`/v1/images/generations`. The audiovisual notebook uses the validated video -generation endpoint for text-to-image, text-to-video, image-to-video, -video-to-video, and synchronized audio. Cosmos3 text-to-image is sent as a -one-frame video request, matching the TensorRT-LLM Cosmos3 pipeline; the notebook -sends it as `num_frames=1`, `seconds=1`, and `fps=8` to satisfy the video request -schema while preserving a single generated frame. Image-to-video and -video-to-video upload their reference media as multipart `input_reference`; -TensorRT-LLM classifies the reference by content. Synchronized audio is enabled -with `enable_audio: true` in `extra_params` and is muxed into the output video. -Keep `ffmpeg` on the server `PATH`: without it, TensorRT-LLM falls back to a -video-only AVI encoder and cannot preserve generated audio. Requests send +The server exposes `/health`, the blocking `/v1/videos/sync`, the asynchronous +`/v1/videos`, and `/v1/images/generations`. The older +`/v1/videos/generations` spelling is a deprecated alias of `/v1/videos/sync`. +The base audiovisual notebook uses `/v1/images/generations` for text-to-image +and `/v1/videos/sync` for text-to-video, image-to-video, video-to-video, and +synchronized audio. Text-to-image sets `extra_params.output_type="image"` and +returns a base64-encoded PNG. Image-to-video uploads multipart +`image_reference`; video-to-video uses `video_reference`. Synchronized audio is +enabled with `enable_audio: true` in `extra_params` and is muxed into the output +video. +Every video request explicitly selects MP4. Keep `ffmpeg` on the server `PATH`: +without it, the request fails early instead of returning browser-incompatible +AVI or dropping generated audio. Requests send Cosmos3 controls through `extra_params`, so use a TensorRT-LLM build that includes the Cosmos3 VisualGen API schema. The notebook sets request-level `max_sequence_length=4096` for longer structured JSON prompts. +Transfer uses the synchronous `/v1/videos/sync` route. For server-derived edge +or blur, upload the raw source video as multipart `video_reference` and set the +corresponding `extra_params` hint to `true`. For a precomputed edge, blur, +depth, segmentation, or WSM control, base64-encode the control inside its hint; +no top-level reference is needed. The server decodes inline media to bytes at the +HTTP boundary. TensorRT-LLM uses `use_guardrails` for its per-request safety +switch; `guardrails`, `control_path`, and other vLLM-Omni-only names are not +interchangeable. + +Action requests use the same synchronous route and upload an image as +`image_reference` or a video as `video_reference`. For the checked-in AV +examples, use the Cosmos Framework reference prompt shown in the Action +cookbook; current TensorRT-LLM ignores the legacy `view_point` field. +Because an action trajectory cannot be represented in MP4 or +AVI, `format=auto` resolves to `safetensors`; the payload contains named `video`, +`action`, and `frame_rate` tensors. The asynchronous `/v1/videos` route also +supports this payload: poll `GET /v1/videos/{id}`, then download it from +`GET /v1/videos/{id}/content`. + ## TensorRT-LLM Reasoner OpenAI-compatible **reasoning** server for image and video understanding. Run diff --git a/cookbooks/cosmos3/generator/action/README.md b/cookbooks/cosmos3/generator/action/README.md index 51ea8268..8cd94264 100644 --- a/cookbooks/cosmos3/generator/action/README.md +++ b/cookbooks/cosmos3/generator/action/README.md @@ -18,15 +18,18 @@ The rest of this doc shows how to run these modes on selected embodiments direct - [Run with Diffusers](#run-with-diffusers) - [Quickstart](#quickstart-1) - [Notebook walkthrough](#diffusers-notebook-walkthrough) -- [Run with vLLM-Omni](#run-with-vllm-omni) +- [Run with TensorRT-LLM](#run-with-tensorrt-llm) - [Quickstart](#quickstart-2) + - [Notebook walkthrough](#tensorrt-llm-notebook-walkthrough) +- [Run with vLLM-Omni](#run-with-vllm-omni) + - [Quickstart](#quickstart-3) - [Notebook walkthrough](#vllm-omni-notebook-walkthrough) - [Post-Train for Cosmos3-Nano-Policy-DROID](#post-train-for-cosmos3-nano-policy-droid) ## Overview -All examples are shown across three different inference backends — native -PyTorch (Cosmos Framework), Diffusers, and vLLM-Omni. Every backend uses the sample +Examples are shown across native PyTorch (Cosmos Framework), Diffusers, +TensorRT-LLM, and vLLM-Omni. Every backend uses the sample assets under [`assets/`](./assets) and covers three tasks: Environment setup for all backends is centralized in the shared @@ -36,7 +39,8 @@ links to the section you need. Generator requires the Guardrail. Request access to the gated [nvidia/Cosmos-1.0-Guardrail](https://huggingface.co/nvidia/Cosmos-1.0-Guardrail) HF repository before running these examples. To disable the guardrail, set -`enable_safety_checker=False` (Diffusers), `guardrails: false` (vLLM-Omni +`enable_safety_checker=False` (Diffusers), `use_guardrails: false` +(TensorRT-LLM `extra_params`), `guardrails: false` (vLLM-Omni `extra_params`/`extra_args`), or `--no-guardrails` (Cosmos Framework). ## Action Definition @@ -165,6 +169,86 @@ before running them. The DROID reader resolves its feature layout from the name it is given, so also set `COSMOS3_DROID_ROOT` to a directory named after the DROID release you have available. +## Run with TensorRT-LLM + +### Quickstart + +Set up and launch the VisualGen server with the one-GPU Nano config: +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), including the +[server-side NLTK data setup](../../README.md#server-side-nltk-data-setup). +Keep `NLTK_DATA` exported in the server's shell. Action requests +upload a conditioning image as multipart `image_reference` or a video as +`video_reference`, and put the mode-specific fields under `extra_params`. +Use the same trained AV prompt as the Cosmos Framework reference; the legacy +`view_point` field is ignored by current TensorRT-LLM. + +Both notebooks' current multipart fields were exercised against TensorRT-LLM +source revision `bca6761ab84fbcd58fc7f914eade7de48b32e35e`. Follow the source +pin in the shared setup. + +```python +import json +from pathlib import Path + +import requests +from safetensors.torch import load as load_safetensors + +action_root = Path("cookbooks/cosmos3/generator/action") +image_path = action_root / "assets/images/av_0.jpg" +actions = json.loads((action_root / "assets/actions/av_traj_forward.json").read_text()) +prompt = "You are an autonomous vehicle planning system." + +with image_path.open("rb") as image_file: + response = requests.post( + "http://localhost:8000/v1/videos/sync", + data={ + "prompt": prompt, + "format": "safetensors", + "response_format": "file", + "seed": "0", + "extra_params": json.dumps( + { + "action_mode": "forward_dynamics", + "domain_name": "av", + "action": actions, + "use_guardrails": True, + } + ), + }, + files={"image_reference": (image_path.name, image_file, "image/jpeg")}, + headers={"Accept": "application/octet-stream"}, + ) +response.raise_for_status() +payload = load_safetensors(response.content) +print(payload["video"].shape, payload["action"].shape, payload["frame_rate"].item()) +``` + +The `av` domain preset supplies its 60-step, 9D, 480p/10 fps recipe; other +recognized domains similarly fill omitted action width, chunk, resolution, and +frame rate. Inverse dynamics predicts its action tensor, while forward dynamics +carries the conditioned action alongside the rollout. Both supported modes use +a tensor output: `format=auto` selects +`safetensors`, and explicit `mp4`/`avi` is rejected rather than dropping the +trajectory. + +The quickstart uses the blocking `POST /v1/videos/sync` route with +`response_format=file`, which returns the tensor payload bytes directly. The +older `/v1/videos/generations` spelling is a deprecated alias. The asynchronous +`POST /v1/videos` route returns HTTP 202; poll +`GET /v1/videos/{id}` and download the same tensor payload from +`GET /v1/videos/{id}/content`. + +### TensorRT-LLM notebook walkthrough + +- [`run_fd_with_trt_llm.ipynb`](./run_fd_with_trt_llm.ipynb) — forward + dynamics from the checked-in AV image and 60×9 action trajectory. +- [`run_id_with_trt_llm.ipynb`](./run_id_with_trt_llm.ipynb) — inverse + dynamics from the checked-in AV observation clip. + +Both notebooks decode TensorRT-LLM's `safetensors` response, write the rollout +video, and inspect the named `action` tensor. This response contract is not the +top-level action JSON used by vLLM-Omni. + ## Run with vLLM-Omni ### Quickstart diff --git a/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb new file mode 100644 index 00000000..0be2bbf9 --- /dev/null +++ b/cookbooks/cosmos3/generator/action/run_fd_with_trt_llm.ipynb @@ -0,0 +1,288 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "fd-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Action Forward Dynamics with TensorRT-LLM\n", + "\n", + "Run the checked-in autonomous-driving example: a first frame plus a 60-step, 9D action\n", + "trajectory conditions a 61-frame rollout. The `av` domain preset supplies the trained 480p,\n", + "10 fps recipe, so the request does not duplicate those mode defaults.\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", + "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install client-side decoding dependencies in this notebook kernel if needed:\n", + "\n", + "```bash\n", + "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", + "```\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-contract", + "metadata": {}, + "source": [ + "## TensorRT-LLM response contract\n", + "\n", + "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", + "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", + "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", + "\n", + "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", + "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", + "`GET /v1/videos/{id}/content`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import json\n", + "import mimetypes\n", + "import time\n", + "\n", + "import imageio.v3 as iio\n", + "import requests\n", + "from IPython.display import Video, display\n", + "from safetensors.torch import load as load_safetensors\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/sync\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "def submit_action(\n", + " prompt: str, reference_path: Path, extra_params: dict, output_path: Path, *, seed: int = 0\n", + ") -> dict:\n", + " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", + " reference_path = reference_path.resolve()\n", + " if not reference_path.exists():\n", + " raise FileNotFoundError(reference_path)\n", + " output_path.parent.mkdir(parents=True, exist_ok=True)\n", + " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", + " form = {\n", + " \"prompt\": prompt,\n", + " \"format\": \"safetensors\",\n", + " \"response_format\": \"file\",\n", + " \"seed\": str(seed),\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"application/octet-stream\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with reference_path.open(\"rb\") as reference_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"image_reference\": (reference_path.name, reference_file, media_type)},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\")\n", + " if \"application/octet-stream\" not in content_type:\n", + " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", + " output_path.write_bytes(response.content)\n", + " payload = load_safetensors(response.content)\n", + " missing = {\"video\", \"action\"} - payload.keys()\n", + " if missing:\n", + " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", + " print(\"saved tensor payload:\", output_path)\n", + " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", + " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", + " return payload\n", + "\n", + "\n", + "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", + " frames = payload[\"video\"].detach().cpu().numpy()\n", + " if frames.ndim == 5 and frames.shape[0] == 1:\n", + " frames = frames[0]\n", + " if frames.ndim != 4 or frames.shape[-1] != 3:\n", + " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", + " fps_tensor = payload.get(\"frame_rate\")\n", + " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", + " iio.imwrite(output_path, frames, fps=fps)\n", + " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", + " display(Video(str(output_path), embed=True))\n", + " return output_path\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-input-title", + "metadata": {}, + "source": [ + "## Prepare the checked-in AV inputs\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-input", + "metadata": {}, + "outputs": [], + "source": [ + "image_path = ACTION_ROOT / \"assets\" / \"images\" / \"av_0.jpg\"\n", + "action_path = ACTION_ROOT / \"assets\" / \"actions\" / \"av_traj_forward.json\"\n", + "actions = json.loads(action_path.read_text())\n", + "assert len(actions) == 60 and all(len(row) == 9 for row in actions)\n", + "\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", + "extra_params = {\n", + " \"action_mode\": \"forward_dynamics\",\n", + " \"domain_name\": \"av\",\n", + " \"action\": actions,\n", + " \"use_guardrails\": True,\n", + "}\n", + "print(\"image:\", image_path)\n", + "print(\"action trajectory:\", action_path, len(actions), \"x\", len(actions[0]))\n", + "print(json.dumps({**extra_params, \"action\": \"<60x9 checked-in trajectory>\"}, indent=2))\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-run-title", + "metadata": {}, + "source": [ + "## Run forward dynamics\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "fd_payload = submit_action(\n", + " prompt,\n", + " image_path,\n", + " extra_params,\n", + " OUTPUT_ROOT / \"forward_dynamics_av.safetensors\",\n", + " seed=0,\n", + ")\n" + ] + }, + { + "cell_type": "markdown", + "id": "fd-trt-view-title", + "metadata": {}, + "source": [ + "## Inspect the rollout and action tensor\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "fd-trt-view", + "metadata": {}, + "outputs": [], + "source": [ + "save_and_view_video(fd_payload, OUTPUT_ROOT / \"forward_dynamics_av.mp4\")\n", + "fd_action = fd_payload[\"action\"].detach().cpu().numpy()\n", + "print(\"first five action rows:\")\n", + "print(fd_action[:5])\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb new file mode 100644 index 00000000..347aa9a7 --- /dev/null +++ b/cookbooks/cosmos3/generator/action/run_id_with_trt_llm.ipynb @@ -0,0 +1,283 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "id-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Action Inverse Dynamics with TensorRT-LLM\n", + "\n", + "Recover an autonomous-driving action trajectory from the checked-in observation clip. The\n", + "`av` domain preset selects the 60-step, 9D trajectory shape and the trained 480p/10 fps recipe.\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then run this\n", + "from the TensorRT-LLM checkout. This notebook is a client; keep the server in a separate terminal.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install client-side decoding dependencies in this notebook kernel if needed:\n", + "\n", + "```bash\n", + "pip install requests safetensors torch imageio imageio-ffmpeg pillow\n", + "```\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "ACTION_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"action\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_ACTION_OUTPUT_ROOT\", ACTION_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"ACTION_ROOT:\", ACTION_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-contract", + "metadata": {}, + "source": [ + "## TensorRT-LLM response contract\n", + "\n", + "The notebooks use the blocking `POST /v1/videos/sync` route. Every Action request sets\n", + "`format=safetensors` and `response_format=file`; the response is `application/octet-stream` with named `video`, `action`,\n", + "and `frame_rate` tensors. This differs from vLLM-Omni's top-level action metadata.\n", + "\n", + "TensorRT-LLM also implements the asynchronous `POST /v1/videos` route. It returns HTTP 202,\n", + "can be polled through `GET /v1/videos/{id}`, and serves the same tensor payload from\n", + "`GET /v1/videos/{id}/content`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import json\n", + "import mimetypes\n", + "import time\n", + "\n", + "import imageio.v3 as iio\n", + "import requests\n", + "from IPython.display import Video, display\n", + "from safetensors.torch import load as load_safetensors\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/sync\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "def submit_action(prompt: str, reference_path: Path, extra_params: dict, output_path: Path) -> dict:\n", + " \"\"\"Call the synchronous API and decode its action/video tensor payload.\"\"\"\n", + " reference_path = reference_path.resolve()\n", + " if not reference_path.exists():\n", + " raise FileNotFoundError(reference_path)\n", + " output_path.parent.mkdir(parents=True, exist_ok=True)\n", + " media_type = mimetypes.guess_type(reference_path.name)[0] or \"application/octet-stream\"\n", + " form = {\n", + " \"prompt\": prompt,\n", + " \"format\": \"safetensors\",\n", + " \"response_format\": \"file\",\n", + " \"seed\": \"0\",\n", + " \"extra_params\": json.dumps(extra_params, separators=(\",\", \":\")),\n", + " }\n", + " headers = {\"Accept\": \"application/octet-stream\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " with reference_path.open(\"rb\") as reference_file:\n", + " response = requests.post(\n", + " video_api_url(),\n", + " data=form,\n", + " files={\"video_reference\": (reference_path.name, reference_file, media_type)},\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\")\n", + " if \"application/octet-stream\" not in content_type:\n", + " raise RuntimeError(f\"Expected a tensor payload, got {content_type!r}\")\n", + " output_path.write_bytes(response.content)\n", + " payload = load_safetensors(response.content)\n", + " missing = {\"video\", \"action\"} - payload.keys()\n", + " if missing:\n", + " raise RuntimeError(f\"TensorRT-LLM action payload is missing {sorted(missing)}; got {sorted(payload)}\")\n", + " print(\"saved tensor payload:\", output_path)\n", + " print(\"video:\", tuple(payload[\"video\"].shape), payload[\"video\"].dtype)\n", + " print(\"action:\", tuple(payload[\"action\"].shape), payload[\"action\"].dtype)\n", + " return payload\n", + "\n", + "\n", + "def save_and_view_video(payload: dict, output_path: Path) -> Path:\n", + " frames = payload[\"video\"].detach().cpu().numpy()\n", + " if frames.ndim == 5 and frames.shape[0] == 1:\n", + " frames = frames[0]\n", + " if frames.ndim != 4 or frames.shape[-1] != 3:\n", + " raise ValueError(f\"Expected video [T,H,W,3], got {frames.shape}\")\n", + " fps_tensor = payload.get(\"frame_rate\")\n", + " fps = float(fps_tensor.item()) if fps_tensor is not None else 24.0\n", + " iio.imwrite(output_path, frames, fps=fps)\n", + " print(\"saved rollout:\", output_path, \"fps:\", fps)\n", + " display(Video(str(output_path), embed=True))\n", + " return output_path\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-input-title", + "metadata": {}, + "source": [ + "## Prepare the checked-in AV observation\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-input", + "metadata": {}, + "outputs": [], + "source": [ + "video_path = ACTION_ROOT / \"assets\" / \"videos\" / \"av_0.mp4\"\n", + "if not video_path.exists():\n", + " raise FileNotFoundError(video_path)\n", + "\n", + "prompt = \"You are an autonomous vehicle planning system.\"\n", + "extra_params = {\n", + " \"action_mode\": \"inverse_dynamics\",\n", + " \"domain_name\": \"av\",\n", + " \"use_guardrails\": True,\n", + "}\n", + "print(\"observation:\", video_path)\n", + "print(json.dumps(extra_params, indent=2))\n", + "display(Video(str(video_path), embed=True))\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-run-title", + "metadata": {}, + "source": [ + "## Run inverse dynamics\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "id_payload = submit_action(\n", + " prompt,\n", + " video_path,\n", + " extra_params,\n", + " OUTPUT_ROOT / \"inverse_dynamics_av.safetensors\",\n", + ")\n" + ] + }, + { + "cell_type": "markdown", + "id": "id-trt-view-title", + "metadata": {}, + "source": [ + "## Inspect the reconstructed video and predicted trajectory\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "id-trt-view", + "metadata": {}, + "outputs": [], + "source": [ + "save_and_view_video(id_payload, OUTPUT_ROOT / \"inverse_dynamics_av.mp4\")\n", + "id_action = id_payload[\"action\"].detach().cpu().numpy()\n", + "print(\"predicted action shape:\", id_action.shape)\n", + "print(\"first five predicted action rows:\")\n", + "print(id_action[:5])\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/audiovisual/README.md b/cookbooks/cosmos3/generator/audiovisual/README.md index f5476b38..627275a1 100644 --- a/cookbooks/cosmos3/generator/audiovisual/README.md +++ b/cookbooks/cosmos3/generator/audiovisual/README.md @@ -272,7 +272,7 @@ prompt = json.load(open("assets/prompts/text2video/robot_kitchen.json")) negative = json.load(open("assets/negative_prompts/text2video/neg_prompt.json")) response = requests.post( - "http://localhost:8000/v1/videos/generations", + "http://localhost:8000/v1/videos/sync", json={ "prompt": json.dumps(prompt, ensure_ascii=True, separators=(",", ":")), "negative_prompt": json.dumps(negative, ensure_ascii=True, separators=(",", ":")), @@ -284,6 +284,8 @@ response = requests.post( "guidance_scale": 6.0, "max_sequence_length": 4096, "seed": 0, + "format": "mp4", + "response_format": "file", "extra_params": { "use_resolution_template": False, "use_duration_template": False, @@ -291,20 +293,26 @@ response = requests.post( "use_guardrails": True, }, }, + headers={"Accept": "video/mp4"}, ) response.raise_for_status() -suffix = ".avi" if "x-msvideo" in response.headers.get("content-type", "") else ".mp4" -Path(f"/tmp/cosmos3_t2v_trtllm{suffix}").write_bytes(response.content) +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_t2v_trtllm.mp4").write_bytes(response.content) ``` For image-to-video, post multipart form data to the same endpoint with the -reference image under `input_reference`. To generate synchronized audio for a +reference image under `image_reference`. To generate synchronized audio for a text-to-video or image-to-video request, add `"enable_audio": True` to `extra_params`. Keep `ffmpeg` installed in the server environment so TensorRT-LLM -can mux the generated audio into the MP4; its fallback AVI encoder is video-only. +can mux the generated audio into MP4. Explicit `format=mp4` makes a missing +encoder fail early instead of returning browser-incompatible AVI. -For video-to-video, upload an MP4 or AVI reference under `input_reference`. -TensorRT-LLM classifies the upload by content and forwards the encoded bytes to +For video-to-video, upload an MP4 reference under `video_reference`. +TensorRT-LLM forwards the encoded bytes to the Cosmos3 workers, which decode the conditioning window with NVDEC: ```python @@ -319,7 +327,7 @@ v2v_negative = json.load(open("assets/negative_prompts/image2video/neg_prompt.js with source_video.open("rb") as video_file: response = requests.post( - "http://localhost:8000/v1/videos/generations", + "http://localhost:8000/v1/videos/sync", data={ "prompt": json.dumps(v2v_prompt, ensure_ascii=True, separators=(",", ":")), "negative_prompt": json.dumps(v2v_negative, ensure_ascii=True, separators=(",", ":")), @@ -330,6 +338,8 @@ with source_video.open("rb") as video_file: "guidance_scale": "6.0", "max_sequence_length": "4096", "seed": "0", + "format": "mp4", + "response_format": "file", "extra_params": json.dumps( { "use_resolution_template": False, @@ -342,12 +352,16 @@ with source_video.open("rb") as video_file: separators=(",", ":"), ), }, - files={"input_reference": (source_video.name, video_file, "video/mp4")}, - headers={"Accept": "video/mp4, video/x-msvideo"}, + files={"video_reference": (source_video.name, video_file, "video/mp4")}, + headers={"Accept": "video/mp4"}, ) response.raise_for_status() -suffix = ".avi" if "x-msvideo" in response.headers.get("content-type", "") else ".mp4" -Path(f"/tmp/cosmos3_v2v_trtllm{suffix}").write_bytes(response.content) +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_v2v_trtllm.mp4").write_bytes(response.content) ``` `condition_video_latent_indexes` identifies clean latent frames in the output; @@ -355,10 +369,10 @@ with `[0, 1]`, TensorRT-LLM consumes the first five pixel frames from the input. Set `condition_video_keep` to `"last"` to condition on the corresponding tail window instead. -For text-to-image, use the same video generation endpoint with `num_frames=1`, -`seconds=1`, and `fps=8`; TensorRT-LLM Cosmos3 returns a one-frame video -response for this path. `num_frames` is passed explicitly so the server does not -derive an eight-frame clip from `seconds * fps`. +For text-to-image, send JSON to `/v1/images/generations` with `format=png`, +`response_format=b64_json`, and `extra_params.output_type="image"`, then +base64-decode `data[0].b64_json`. This produces a PNG image instead of wrapping +one frame in a video container. To run **Cosmos3-Edge** instead, serve `nvidia/Cosmos3-Edge` on a single GPU with no config override (`trtllm-serve nvidia/Cosmos3-Edge --port 8000`) and send diff --git a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb index 2111257e..fc0e1cf7 100644 --- a/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb +++ b/cookbooks/cosmos3/generator/audiovisual/run_with_trt_llm.ipynb @@ -29,9 +29,9 @@ "source": [ "## 1. Prerequisites\n", "\n", - "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Generation uses `/v1/videos/generations`; Cosmos3 text-to-image runs as a one-frame VisualGen video request with `num_frames=1`, `seconds=1`, and `fps=8`. Image-to-video and video-to-video send their reference media as multipart `input_reference`.\n", + "Use a running TensorRT-LLM server with Cosmos3 VisualGen audio and video-to-video support, and set endpoint environment variables before the setup cell if you are not using the local default. Text-to-image uses `/v1/images/generations`; video generation uses `/v1/videos/sync`. Image-to-video uploads multipart `image_reference`, while video-to-video uploads multipart `video_reference`.\n", "\n", - "Install the `ffmpeg` CLI in the server environment. TensorRT-LLM uses it to encode MP4 and mux synchronized audio; without it, the fallback AVI encoder drops audio.\n", + "Install the `ffmpeg` CLI in the server environment. This notebook explicitly requests browser-compatible MP4, so TensorRT-LLM returns an actionable error if the encoder is unavailable; `ffmpeg` also muxes synchronized audio.\n", "\n", "Generator requires the Guardrail. Request access to the gated [nvidia/Cosmos-1.0-Guardrail](https://huggingface.co/nvidia/Cosmos-1.0-Guardrail) HF repository before running these examples. TensorRT-LLM loads guardrails by default; to disable them, set `TRTLLM_DISABLE_COSMOS3_GUARDRAILS=1` before starting the server or set `use_guardrails` to `False` in the request `extra_params`.\n", "\n", @@ -60,7 +60,9 @@ "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", "\n", - "trtllm-serve nvidia/Cosmos3-Nano --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" --port 8000\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", "```\n", "\n", "### Cosmos3-Super\n", @@ -70,7 +72,10 @@ "```bash\n", "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", "\n", - "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve nvidia/Cosmos3-Super --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" --port 8000\n", + "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \\\n", + " nvidia/Cosmos3-Super \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" \\\n", + " --port 8000\n", "```\n", "\n", "### Cosmos3-Edge\n", @@ -81,7 +86,7 @@ "trtllm-serve nvidia/Cosmos3-Edge --port 8000\n", "```\n", "\n", - "TensorRT-LLM serves Edge for text-to-image, text-to-video, and image-to-video. Text-to-image goes to `/v1/images/generations` with `output_type=\"image\"`; the two video modes go to `/v1/videos/generations`. Edge has no audio tower, so `enable_audio` is unavailable, and the checkpoint's action weights are not served by this pipeline. Requests outside the model card's validated envelope (256p/480p, 50-150 frames, 12-30 FPS) still run and log an advisory line.\n", + "TensorRT-LLM serves Edge for text-to-image, text-to-video, and image-to-video. Text-to-image goes to `/v1/images/generations` with `output_type=\"image\"`; the two video modes go to `/v1/videos/sync`. Edge has no audio tower, so `enable_audio` is unavailable, and the checkpoint's action weights are not served by this pipeline. Requests outside the model card's validated envelope (256p/480p, 50-150 frames, 12-30 FPS) still run and log an advisory line.\n", "\n", "TensorRT-LLM exposes `/health` when the server is ready. This notebook sends Cosmos3 model-specific controls through `extra_params`, so use a TensorRT-LLM release that includes the Cosmos3 VisualGen API schema.\n" ], @@ -172,7 +177,11 @@ "\n", "\n", "def video_api_url(base_url: str) -> str:\n", - " return f\"{api_root_url(base_url)}/videos/generations\"\n", + " return f\"{api_root_url(base_url)}/videos/sync\"\n", + "\n", + "\n", + "def image_api_url(base_url: str) -> str:\n", + " return f\"{api_root_url(base_url)}/images/generations\"\n", "\n", "\n", "for model, base_url in TRTLLM_ENDPOINTS.items():\n", @@ -180,7 +189,8 @@ " print(model)\n", " print(\" api root:\", api_root_url(base_url))\n", " print(\" health:\", health_url(base_url))\n", - " print(\" videos generations:\", video_api_url(base_url))\n", + " print(\" images:\", image_api_url(base_url))\n", + " print(\" videos sync:\", video_api_url(base_url))\n", " print(\" scheme:\", parsed.scheme)\n", " print(\" host:\", parsed.netloc)\n" ], @@ -281,7 +291,7 @@ "\n", "\n", "def video_api_url(base_url: str) -> str:\n", - " return f\"{api_root_url(base_url)}/videos/generations\"\n", + " return f\"{api_root_url(base_url)}/videos/sync\"\n", "\n", "\n", "def image_api_url(base_url: str) -> str:\n", @@ -334,8 +344,8 @@ " \"use_system_prompt\": False,\n", " \"use_guardrails\": True,\n", "}\n", - "# TensorRT-LLM Cosmos3 uses the stable video endpoint path here. Text-to-image\n", - "# generation is represented as a one-frame video request.\n", + "# TensorRT-LLM uses the image endpoint for T2I and the synchronous video\n", + "# endpoint for T2V, I2V, and V2V.\n", "ASSET_SETS = {\n", " \"t2i_nano\": {\n", " \"model\": \"Cosmos3-Nano\",\n", @@ -487,6 +497,8 @@ " extra_params = dict(TRTLLM_EXTRA_PARAMS)\n", " extra_params[\"enable_audio\"] = bool(spec.get(\"enable_audio\", False))\n", " extra_params.update(spec.get(\"extra_params\", {}))\n", + " if spec[\"mode\"] == \"text2image\":\n", + " extra_params[\"output_type\"] = \"image\"\n", " payload = {\n", " \"model_mode\": spec[\"mode\"],\n", " \"name\": use_case,\n", @@ -495,11 +507,8 @@ " **spec.get(\"sampling\", FIXED_SAMPLING),\n", " }\n", " if spec[\"mode\"] == \"text2image\":\n", - " if extra_params.get(\"output_type\") != \"image\":\n", - " # Legacy path: text-to-image expressed as a one-frame video request.\n", - " payload[\"num_frames\"] = 1\n", - " payload[\"fps\"] = 8\n", - " payload[\"seconds\"] = 1\n", + " payload.pop(\"num_frames\", None)\n", + " payload.pop(\"fps\", None)\n", " else:\n", " negative_prompt_path = asset_path(\n", " spec.get(\"negative_prompt\") or f\"assets/negative_prompts/{spec['mode']}/neg_prompt.json\"\n", @@ -522,10 +531,7 @@ " print(f\"output: {output_dir}\")\n", " print(f\"prompt: {prompt_path.relative_to(COSMOS_ROOT)}\")\n", " if spec[\"mode\"] == \"text2image\":\n", - " if extra_params.get(\"output_type\") == \"image\":\n", - " print(\"note: native image mode; the server's default negative prompt is empty\")\n", - " else:\n", - " print(\"note: TensorRT-LLM Cosmos3 text-to-image is served as a one-frame video response\")\n", + " print(\"note: TensorRT-LLM Cosmos3 text-to-image uses /v1/images/generations\")\n", " if \"vision_path\" in payload:\n", " image_display_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", " print(f\"image: {image_display_path.relative_to(COSMOS_ROOT)}\")\n", @@ -536,14 +542,12 @@ " print(\"note: condition_video_latent_indexes=[0, 1] consumes five input pixel frames\")\n", " if payload[\"extra_params\"][\"enable_audio\"]:\n", " print(\"audio: enabled; ffmpeg is required on the server to mux it into MP4\")\n", - " preview_keys = [\n", - " key\n", - " for key in [\n", - " \"model_mode\", \"name\", \"num_steps\", \"guidance\", \"seconds\", \"fps\", \"num_frames\",\n", - " \"max_sequence_length\", \"resolution\", \"aspect_ratio\", \"seed\", \"extra_params\",\n", - " ]\n", - " if key in payload\n", - " ]\n", + " preview_keys = [\"model_mode\", \"name\", \"num_steps\", \"guidance\"]\n", + " if spec[\"mode\"] != \"text2image\":\n", + " preview_keys.extend([\"fps\", \"num_frames\"])\n", + " preview_keys.extend([\"max_sequence_length\", \"resolution\", \"aspect_ratio\", \"seed\", \"extra_params\"])\n", + " if \"seconds\" in payload:\n", + " preview_keys.insert(5, \"seconds\")\n", " print(json.dumps({k: payload[k] for k in preview_keys}, indent=2))\n", " return payload_path, output_dir, spec[\"model\"]\n", "\n", @@ -575,11 +579,15 @@ " \"seconds\": payload.get(\"seconds\", payload[\"num_frames\"] / payload[\"fps\"]),\n", " \"fps\": payload[\"fps\"],\n", " \"num_frames\": payload[\"num_frames\"],\n", - " \"num_inference_steps\": payload[\"num_steps\"],\n", - " \"guidance_scale\": payload[\"guidance\"],\n", " \"max_sequence_length\": payload[\"max_sequence_length\"],\n", " \"seed\": payload[\"seed\"],\n", + " \"format\": \"mp4\",\n", + " \"response_format\": \"file\",\n", " }\n", + " if payload.get(\"num_steps\") is not None:\n", + " body[\"num_inference_steps\"] = payload[\"num_steps\"]\n", + " if payload.get(\"guidance\") is not None:\n", + " body[\"guidance_scale\"] = payload[\"guidance\"]\n", " if payload.get(\"negative_prompt\") is not None:\n", " body[\"negative_prompt\"] = payload[\"negative_prompt\"]\n", " body[\"extra_params\"] = payload[\"extra_params\"]\n", @@ -587,24 +595,23 @@ "\n", "\n", "def build_trtllm_image_body(payload: dict) -> dict:\n", - " \"\"\"Native text-to-image request for the image generation endpoint.\n", - "\n", - " The images API carries no frame or frame-rate fields. ``output_type=\"image\"``\n", - " in ``extra_params`` is what selects the image path; without it the server runs\n", - " video mode, which also swaps in Cosmos3's video negative prompt by default.\n", - " \"\"\"\n", " height, width = payload_dimensions(payload)\n", - " return {\n", + " body = {\n", " \"prompt\": payload[\"prompt\"],\n", " \"size\": f\"{width}x{height}\",\n", - " \"num_inference_steps\": payload[\"num_steps\"],\n", - " \"guidance_scale\": payload[\"guidance\"],\n", + " \"format\": \"png\",\n", + " \"response_format\": \"b64_json\",\n", " \"max_sequence_length\": payload[\"max_sequence_length\"],\n", " \"seed\": payload[\"seed\"],\n", - " \"response_format\": \"b64_json\",\n", - " \"format\": \"png\",\n", " \"extra_params\": payload[\"extra_params\"],\n", " }\n", + " if payload.get(\"num_steps\") is not None:\n", + " body[\"num_inference_steps\"] = payload[\"num_steps\"]\n", + " if payload.get(\"guidance\") is not None:\n", + " body[\"guidance_scale\"] = payload[\"guidance\"]\n", + " if payload.get(\"negative_prompt\") is not None:\n", + " body[\"negative_prompt\"] = payload[\"negative_prompt\"]\n", + " return body\n", "\n", "\n", "def _auth_headers() -> list[str]:\n", @@ -612,16 +619,12 @@ " return [\"-H\", f\"Authorization: Bearer {api_key}\"] if api_key else []\n", "\n", "\n", - "def _video_extension_from_response(header_text: str, content_path: Path) -> str:\n", - " lowered = header_text.lower()\n", - " if \"video/x-msvideo\" in lowered or \"video/avi\" in lowered:\n", - " return \".avi\"\n", + "def _require_mp4_response(header_text: str, content_path: Path) -> None:\n", + " if \"video/mp4\" not in header_text.lower():\n", + " raise RuntimeError(f\"Expected video/mp4 response, got headers:\\n{header_text}\")\n", " head = content_path.read_bytes()[:16]\n", - " if head.startswith(b\"RIFF\") and b\"AVI\" in head[:12]:\n", - " return \".avi\"\n", - " if len(head) >= 12 and head[4:8] == b\"ftyp\":\n", - " return \".mp4\"\n", - " return \".mp4\"\n", + " if len(head) < 12 or head[4:8] != b\"ftyp\":\n", + " raise RuntimeError(\"TensorRT-LLM returned a non-MP4 payload\")\n", "\n", "\n", "def post_video(*, payload_path: Path, payload: dict, output_stem: Path, model: str) -> Path:\n", @@ -644,7 +647,7 @@ " \"-D\",\n", " str(header_path),\n", " \"-H\",\n", - " \"Accept: video/mp4, video/x-msvideo, application/octet-stream\",\n", + " \"Accept: video/mp4\",\n", " ]\n", " cmd += _auth_headers()\n", "\n", @@ -657,10 +660,12 @@ " if payload[\"model_mode\"] == \"image2video\":\n", " reference_path = resolve_payload_path(payload_path, payload[\"vision_path\"])\n", " reference_type = \"image/jpeg\"\n", + " reference_field = \"image_reference\"\n", " else:\n", " reference_path = resolve_payload_path(payload_path, payload[\"video_path\"])\n", " reference_type = \"video/mp4\"\n", - " cmd += [\"-F\", f\"input_reference=@{reference_path};type={reference_type}\"]\n", + " reference_field = \"video_reference\"\n", + " cmd += [\"-F\", f\"{reference_field}=@{reference_path};type={reference_type}\"]\n", " else:\n", " cmd += [\"-H\", \"Content-Type: application/json\"]\n", " cmd += [\"-d\", json.dumps(body, separators=(\",\", \":\"))]\n", @@ -672,8 +677,8 @@ " raise RuntimeError(f\"TensorRT-LLM request failed with exit code {result.returncode}; see {error_path}\")\n", "\n", " header_text = header_path.read_text() if header_path.exists() else \"\"\n", - " ext = _video_extension_from_response(header_text, tmp_path)\n", - " output_path = output_stem.with_suffix(ext)\n", + " _require_mp4_response(header_text, tmp_path)\n", + " output_path = output_stem.with_suffix(\".mp4\")\n", " if output_path.exists():\n", " output_path.unlink()\n", " tmp_path.replace(output_path)\n", @@ -682,11 +687,12 @@ "\n", "def post_image(*, payload: dict, output_stem: Path, model: str) -> Path:\n", " url = image_api_url(TRTLLM_ENDPOINTS[model])\n", + " response_path = Path(f\"{output_stem}.response.json\")\n", " error_path = Path(f\"{output_stem}.error.txt\")\n", - " if error_path.exists():\n", - " error_path.unlink()\n", + " for path in [response_path, error_path]:\n", + " if path.exists():\n", + " path.unlink()\n", "\n", - " body = build_trtllm_image_body(payload)\n", " cmd = [\n", " \"curl\",\n", " \"-sS\",\n", @@ -695,25 +701,28 @@ " \"POST\",\n", " url,\n", " \"-H\",\n", - " \"Accept: application/json\",\n", - " \"-H\",\n", " \"Content-Type: application/json\",\n", + " \"-d\",\n", + " json.dumps(build_trtllm_image_body(payload), separators=(\",\", \":\")),\n", + " \"-o\",\n", + " str(response_path),\n", " ]\n", " cmd += _auth_headers()\n", - " cmd += [\"-d\", json.dumps(body, separators=(\",\", \":\"))]\n", - "\n", " result = subprocess.run(cmd, text=True, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n", " if result.returncode != 0:\n", - " error_path.write_text((result.stdout or \"\") + (result.stderr or \"\"))\n", - " raise RuntimeError(\n", - " f\"TensorRT-LLM image request failed with exit code {result.returncode}; see {error_path}\"\n", - " )\n", + " server_body = response_path.read_text() if response_path.exists() else \"\"\n", + " error_path.write_text(server_body + (result.stdout or \"\") + (result.stderr or \"\"))\n", + " raise RuntimeError(f\"TensorRT-LLM request failed with exit code {result.returncode}; see {error_path}\")\n", "\n", - " response = json.loads(result.stdout)\n", - " output_path = output_stem.with_suffix(f\".{response.get('output_format', 'png')}\")\n", - " if output_path.exists():\n", - " output_path.unlink()\n", - " output_path.write_bytes(base64.b64decode(response[\"data\"][0][\"b64_json\"]))\n", + " response = json.loads(response_path.read_text())\n", + " try:\n", + " image_bytes = base64.b64decode(response[\"data\"][0][\"b64_json\"], validate=True)\n", + " except (KeyError, IndexError, TypeError, ValueError) as exc:\n", + " raise RuntimeError(f\"Unexpected TensorRT-LLM image response: {response}\") from exc\n", + " if not image_bytes.startswith(b\"\\x89PNG\\r\\n\\x1a\\n\"):\n", + " raise RuntimeError(\"TensorRT-LLM returned a non-PNG image payload\")\n", + " output_path = output_stem.with_suffix(\".png\")\n", + " output_path.write_bytes(image_bytes)\n", " return output_path\n", "\n", "\n", @@ -723,8 +732,11 @@ " output_dir.mkdir(parents=True, exist_ok=True)\n", " payload = json.loads(payload_path.read_text())\n", " output_stem = output_dir / payload[\"name\"]\n", - " is_image = payload[\"extra_params\"].get(\"output_type\") == \"image\"\n", - " endpoint = (image_api_url if is_image else video_api_url)(TRTLLM_ENDPOINTS[model])\n", + " endpoint = (\n", + " image_api_url(TRTLLM_ENDPOINTS[model])\n", + " if payload[\"model_mode\"] == \"text2image\"\n", + " else video_api_url(TRTLLM_ENDPOINTS[model])\n", + " )\n", " print(\"endpoint:\", endpoint)\n", " print(\"payload:\", payload_path)\n", " print(\"output stem:\", output_stem)\n", @@ -733,19 +745,16 @@ " if payload[\"model_mode\"] == \"video2video\":\n", " print(\"input video:\", resolve_payload_path(payload_path, payload[\"video_path\"]))\n", " t0 = time.time()\n", - " if is_image:\n", + " if payload[\"model_mode\"] == \"text2image\":\n", " output_path = post_image(payload=payload, output_stem=output_stem, model=model)\n", " else:\n", - " output_path = post_video(\n", - " payload_path=payload_path, payload=payload, output_stem=output_stem, model=model\n", - " )\n", + " output_path = post_video(payload_path=payload_path, payload=payload, output_stem=output_stem, model=model)\n", " print(f\"wrote {output_path} in {time.time() - t0:.1f}s\")\n", " return output_path\n", "\n", "\n", "def display_video(path: Path, *, width: int = 720) -> None:\n", - " suffix = path.suffix.lower()\n", - " media_type = \"video/x-msvideo\" if suffix == \".avi\" else \"video/mp4\"\n", + " media_type = \"video/mp4\"\n", " data = base64.b64encode(path.read_bytes()).decode(\"ascii\")\n", " label = html.escape(str(path))\n", " markup = f\"\"\"\n", @@ -759,21 +768,19 @@ "\n", "def view_run(output_dir: str | Path) -> None:\n", " output_dir = Path(output_dir)\n", + " images = [path for path in sorted(output_dir.rglob(\"*.png\"))]\n", " videos = [\n", " path\n", " for path in sorted(output_dir.rglob(\"*\"))\n", - " if path.suffix.lower() in {\".mp4\", \".avi\"}\n", + " if path.suffix.lower() == \".mp4\"\n", " and not path.name.endswith((\"_preview.mp4\", \"_browser.mp4\"))\n", " ]\n", - " images = [\n", - " path for path in sorted(output_dir.rglob(\"*\")) if path.suffix.lower() in IMAGE_EXTENSIONS\n", - " ]\n", - " if not videos and not images:\n", + " if not images and not videos:\n", " print(f\"No generated media found under {output_dir}\")\n", " return\n", " for src in images:\n", " print(f\"source: {src} ({src.stat().st_size // 1024} KB)\")\n", - " display(Image(filename=str(src), width=640))\n", + " display(Image(filename=str(src), width=720))\n", " for src in videos:\n", " print(f\"source: {src} ({src.stat().st_size // 1024} KB)\")\n", " display_video(src)\n" @@ -804,7 +811,7 @@ "source": [ "## Nano: Text to Image\n", "\n", - "Nano text-to-image generation using a structured JSON prompt. TensorRT-LLM Cosmos3 returns this as a one-frame video from the video generation endpoint; the request uses `num_frames=1`, `seconds=1`, and `fps=8`.\n" + "Nano text-to-image generation using a structured JSON prompt. TensorRT-LLM returns a base64-encoded PNG from `/v1/images/generations`.\n" ], "id": "nano-t2i" }, @@ -1032,7 +1039,7 @@ "source": [ "## Nano: Image to Video with Audio\n", "\n", - "Nano image-to-video generation using its paired coastal-road image and synchronized audio prompt. The reference image is uploaded as multipart `input_reference`, and `enable_audio` is set in `extra_params`.\n" + "Nano image-to-video generation using its paired coastal-road image and synchronized audio prompt. The reference image is uploaded as multipart `image_reference`, and `enable_audio` is set in `extra_params`.\n" ], "id": "nano-i2av" }, @@ -1089,7 +1096,7 @@ "source": [ "## Nano: Video to Video\n", "\n", - "Nano video-to-video generation using a checked-in MP4 reference. TensorRT-LLM classifies the multipart `input_reference` by content and forwards the encoded video to each worker for NVDEC decoding. `condition_video_latent_indexes=[0, 1]` pins the first two output latent frames, which consumes the first five pixel frames of the reference.\n" + "Nano video-to-video generation using a checked-in MP4 `video_reference`. TensorRT-LLM forwards the encoded video to each worker for NVDEC decoding. `condition_video_latent_indexes=[0, 1]` pins the first two output latent frames, which consumes the first five pixel frames of the reference.\n" ], "id": "nano-v2v" }, @@ -1156,7 +1163,7 @@ "source": [ "## Super: Text to Image\n", "\n", - "Super text-to-image generation using the same structured JSON prompt. TensorRT-LLM Cosmos3 returns this as a one-frame video from the video generation endpoint; the request uses `num_frames=1`, `seconds=1`, and `fps=8`.\n" + "Super text-to-image generation using the same structured JSON prompt. TensorRT-LLM returns a base64-encoded PNG from `/v1/images/generations`.\n" ], "id": "super-t2i" }, diff --git a/cookbooks/cosmos3/generator/transfer/README.md b/cookbooks/cosmos3/generator/transfer/README.md index e3467e4d..735c776a 100644 --- a/cookbooks/cosmos3/generator/transfer/README.md +++ b/cookbooks/cosmos3/generator/transfer/README.md @@ -2,7 +2,7 @@ Cosmos3 video **transfer** examples — **Nano** (single GPU) and **Super** (multi-GPU, 32B) — on the native PyTorch (Cosmos Framework) path, the Diffusers modular pipeline, and the -OpenAI-compatible vLLM-Omni server path. +OpenAI-compatible TensorRT-LLM and vLLM-Omni server paths. Sample assets under [`assets/`](./assets) cover spatial control signals paired with `prompt.json` files: @@ -13,10 +13,10 @@ Sample assets under [`assets/`](./assets) cover spatial control signals paired w - **World scenario (WSM)** — world-scenario map control plus caption. - **Multi-control** — two or more hints; Cosmos Framework also supports per-hint weights. -Both Cosmos Framework and vLLM-Omni support multi-control transfer. Per-hint -weighting is supported only by Cosmos Framework; vLLM-Omni accepts multiple -controls but does not support per-hint weights. Diffusers accepts multiple -precomputed controls, also without per-hint weights. +Cosmos Framework, TensorRT-LLM, and vLLM-Omni support multi-control transfer. +Per-hint weighting is supported only by Cosmos Framework; the serving backends +accept multiple unweighted controls. Diffusers also accepts multiple +precomputed controls without per-hint weights. Environment setup is centralized in the shared [Cosmos3 cookbooks environment setup](../../README.md) guide. @@ -25,6 +25,9 @@ Environment setup is centralized in the shared Video transfer generates a target clip from a `prompt.json` caption and one or more spatial control signals. The Framework path uses `model_mode` `video2video` in a local JSON spec. +TensorRT-LLM uses `POST /v1/videos/sync`. A raw source used to derive edge or +blur is multipart `video_reference`; precomputed controls are carried directly +under their `extra_params` hint and do not need a top-level reference. The vLLM-Omni path uses `POST /v1/videos/sync` and passes one or more hint keys (`edge`, `blur`, `depth`, `seg`, or `wsm`) inside `extra_params`. Cosmos Framework accepts pre-computed control videos (`control_path`) or derives active controls from a raw source video @@ -40,7 +43,7 @@ for the negative caption. | Blur | `assets/blur/` | `control_blur.mp4` + `prompt.json` | 121 frames @ 30 FPS | | Depth | `assets/depth/` | `control_depth.mp4` + `prompt.json` | 121 frames @ 30 FPS | | Segmentation | `assets/seg/` | `control_seg.mp4` + `prompt.json` | 121 frames @ 30 FPS | -| World scenario (WSM) | `assets/wsm/` | `control_wsm.mp4` + `prompt.json` | 101 frames @ 10 FPS | +| World scenario (WSM) | `assets/wsm/` | `control_wsm.mp4` + `prompt.json` | 100 frames @ 10 FPS | | Multi-control | `assets/multi_control/` | `vision_path` + multiple hints (Framework example) | 121 frames @ 30 FPS | Transfer inference is selected automatically when any hint key is present in the @@ -191,6 +194,96 @@ by the same five controls on Cosmos3-Super. It reads the same [`specs/`](./specs reuses the previews from [`preview_helpers.py`](./preview_helpers.py), writing outputs to `outputs/notebooks/diffusers//transfer_/vision.mp4`. +## Run with TensorRT-LLM + +### Quickstart + +Set up and launch Nano or Super as described in the shared +[TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), including the +[server-side NLTK data setup](../../README.md#server-side-nltk-data-setup). +Keep `NLTK_DATA` exported in the server's shell. TensorRT-LLM +accepts a raw source video as multipart `video_reference` when edge or blur will +be computed on the server. The checked-in assets are already precomputed +controls, so this example base64-encodes the control inside its hint and does +not send a top-level reference. The server decodes inline media at the HTTP +boundary before validating and dispatching the request. + +```python +import base64 +import json +from pathlib import Path + +import requests + +transfer_root = Path("cookbooks/cosmos3/generator/transfer") +control_path = transfer_root / "assets/depth/control_depth.mp4" +prompt = json.dumps(json.load(open(transfer_root / "assets/depth/prompt.json"))) +negative = json.dumps(json.load(open(transfer_root / "assets/negative_prompt.json"))) +extra_params = { + "use_resolution_template": False, + "use_duration_template": False, + "use_system_prompt": False, + "use_guardrails": True, + "flow_shift": 10.0, + "depth": { + "control": base64.b64encode(control_path.read_bytes()).decode("ascii") + }, + "control_guidance": 1.5, + "num_video_frames_per_chunk": 121, + "num_conditional_frames": 1, + "num_first_chunk_conditional_frames": 0, + "share_vision_temporal_positions": True, + "emphasize_control_in_prompt": False, + "max_frames": 121, +} + +response = requests.post( + "http://localhost:8000/v1/videos/sync", + json={ + "prompt": prompt, + "negative_prompt": negative, + "size": "1280x720", + "num_frames": 121, + "fps": 30, + "num_inference_steps": 50, + "guidance_scale": 3.0, + "max_sequence_length": 4096, + "seed": 2026, + "format": "mp4", + "response_format": "file", + "extra_params": extra_params, + }, + headers={"Accept": "video/mp4"}, +) +response.raise_for_status() +if ( + "video/mp4" not in response.headers.get("content-type", "") + or response.content[4:8] != b"ftyp" +): + raise RuntimeError("TensorRT-LLM did not return browser-compatible MP4") +Path("/tmp/cosmos3_transfer_depth_trtllm.mp4").write_bytes(response.content) +``` + +Only edge and blur can be generated from a raw uploaded source (`"edge": true` +or `"blur": true`); depth, segmentation, and WSM always require precomputed +control media. Precomputed edge and blur are also accepted. Use +`use_guardrails`, not vLLM-Omni's `guardrails`, and send encoded control bytes +rather than a server-local `control_path`. When neither output dimension is +specified, TensorRT-LLM chooses the nearest supported bucket from the source or +first precomputed control's aspect ratio. The checked-in blur example explicitly +requests its native 4:3 bucket, 1104×832; the other examples request +1280×720. All preserve their native frame count/fps, and WSM uses 100 frames +at 10 fps. `/v1/videos/generations` remains only as a deprecated alias of the +canonical blocking `/v1/videos/sync` route. + +### TensorRT-LLM notebook walkthrough + +[`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) +runs edge, blur, depth, segmentation, and WSM against an already-running Nano +or Super server. It reuses the checked-in media and captions, shows the complete +mode matrix, validates the JSON/base64 control request contract, and previews the +synchronous encoded-video responses. + ## Run with vLLM-Omni ### Quickstart @@ -315,6 +408,9 @@ Key fields: - [`run_video_transfer_with_diffusers.ipynb`](./run_video_transfer_with_diffusers.ipynb) — full tutorial for the Diffusers modular pipeline: five single-control transfers on Nano, then the same five on Super, driven by the same specs. +- [`run_video_transfer_with_trt_llm.ipynb`](./run_video_transfer_with_trt_llm.ipynb) — + five single-control transfers through TensorRT-LLM's synchronous VisualGen API, + using base64-encoded precomputed controls in JSON `extra_params`. - [`run_video_transfer_with_vllm_omni.ipynb`](./run_video_transfer_with_vllm_omni.ipynb) — full tutorial against an already-running vLLM-Omni server: endpoint checks, repo-local control paths, five single-control transfer requests, and compact previews. The API diff --git a/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json b/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json index 90f2dd01..59ba72f3 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/depth/prompt.json @@ -62,7 +62,7 @@ }, { "description": "A white sedan driving slowly forward in the center of the road", - "appearance_details": "Clean white paint, standard compact sedan body, brake lights occasionally glowing red", + "appearance_details": "Clean white paint, standard compact sedan body, rear lights occasionally glowing red", "relationship": "Directly ahead of the observer, pacing the forward motion", "location": "Center foreground to midground", "relative_size": "Medium within frame", @@ -198,4 +198,4 @@ "aspect_ratio": "16,9", "duration": "4s", "fps": 30 -} \ No newline at end of file +} diff --git a/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json b/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json index ec91950b..efdaee76 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/edge/prompt.json @@ -21,7 +21,7 @@ "number_of_legs": 2 } ], - "background_setting": "A brightly lit rehearsal studio with light-colored wood-style flooring marked by circular green decals used as position markers. One wall is covered by a large black curtain decorated with scattered small white rectangular papers, in front of which two tripods hold small recording cameras. The opposite wall features a large window with bright red frames through which natural daylight streams, revealing a brick building across the way. Beneath the window rests a long bench with alternating orange and white patterned cushions, and a black floor-standing speaker stands nearby.", + "background_setting": "A brightly lit rehearsal studio with light-colored wood-style flooring marked by circular green decals used as position markers. One wall is covered by a large black curtain decorated with scattered small white rectangular papers, in front of which two tripods hold small recording cameras. The opposite wall features a large window with bright red frames through which natural daylight streams, revealing a masonry building across the way. Beneath the window rests a long bench with alternating orange and white patterned cushions, and a black floor-standing speaker stands nearby.", "lighting": { "conditions": "Bright daylight combined with even studio lighting", "direction": "Side-lit from the right through the large red-framed window, with soft ambient fill across the room", @@ -83,4 +83,4 @@ "aspect_ratio": "16,9", "duration": "4s", "fps": 30 -} \ No newline at end of file +} diff --git a/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json b/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json index 44dff269..0b34964d 100644 --- a/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json +++ b/cookbooks/cosmos3/generator/transfer/assets/negative_prompt.json @@ -2,7 +2,7 @@ "subjects": [ { "description": "Blurry, poorly defined subjects with inconsistent shapes and unrealistic proportions.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -22,7 +22,7 @@ }, { "description": "Extremely low-quality subjects with visible rendering artifacts, broken mesh geometry, and completely unrealistic proportions throughout.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -42,7 +42,7 @@ }, { "description": "Poorly generated subjects exhibiting all hallmarks of failed neural rendering \u2014 flickering edges, inconsistent depth, and uncanny spatial relationships.", - "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, color bleeding between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", + "appearance_details": "Distorted features, visible compression artifacts, muddy textures lacking fine detail, colors improperly mixing between elements, and unnatural skin tones or surface textures that appear artificial or computer-generated.", "relationship": "Subjects appear disconnected from the environment, floating or improperly grounded in the scene without proper occlusion or spatial coherence.", "location": "Subjects are poorly placed within the frame, appearing at awkward positions that violate basic compositional rules.", "relative_size": "Inconsistent scale relationships between subjects and the environment, with objects appearing too large or too small relative to their surroundings.", @@ -71,7 +71,7 @@ "aesthetics": { "composition": "Cluttered, poorly framed composition with no clear focal point. Important elements are cut off by the frame edges. The rule of thirds is completely ignored, leading to an unbalanced and visually unpleasant arrangement.", "color_scheme": "Oversaturated, garish colors that clash violently. Color banding is visible in gradient areas. The overall palette feels artificial and digitally processed rather than natural.", - "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels lifeless and sterile despite attempting to portray dynamic action.", + "mood_atmosphere": "Unsettling, uncanny atmosphere that fails to evoke any intended emotional response. The scene feels flat and sterile despite attempting to portray dynamic action.", "patterns": "Visible tiling artifacts in textures, moir\u00e9 patterns, and aliasing on edges." }, "cinematography": { diff --git a/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb new file mode 100644 index 00000000..a5955f17 --- /dev/null +++ b/cookbooks/cosmos3/generator/transfer/run_video_transfer_with_trt_llm.ipynb @@ -0,0 +1,437 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "transfer-trt-license", + "metadata": {}, + "source": [ + "\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-title", + "metadata": {}, + "source": [ + "# Cosmos3 Video Transfer with TensorRT-LLM\n", + "\n", + "Run the five established controls—Canny/edge, blur, depth, segmentation, and WSM—against an\n", + "already-running TensorRT-LLM VisualGen server. The same notebook works with Cosmos3-Nano or\n", + "Cosmos3-Super; select the checkpoint when launching the server.\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-server", + "metadata": {}, + "source": [ + "## Start the TensorRT-LLM server\n", + "\n", + "Follow the shared [TensorRT-LLM setup](../../README.md#tensorrt-llm-generator), then launch one\n", + "server from the TensorRT-LLM checkout.\n", + "Complete the [server-side NLTK data setup](../../README.md#server-side-nltk-data-setup) there first, and keep `NLTK_DATA` exported in that server terminal.\n", + "\n", + "```bash\n", + "export TRTLLM_ROOT=\"${TRTLLM_ROOT:-$PWD}\"\n", + "\n", + "# Nano: one GPU\n", + "trtllm-serve nvidia/Cosmos3-Nano \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-nano-1gpu.yaml\" \\\n", + " --port 8000\n", + "\n", + "# Super: four GPUs (run instead of Nano)\n", + "torchrun --nproc_per_node=4 -m tensorrt_llm.commands.serve \\\n", + " nvidia/Cosmos3-Super \\\n", + " --visual_gen_args \"$TRTLLM_ROOT/examples/visual_gen/configs/cosmos3-super-4gpu.yaml\" \\\n", + " --port 8000\n", + "```\n", + "\n", + "Install `requests` in this notebook kernel if needed, and install `ffmpeg` in the server\n", + "environment. The response is synchronous MP4 video\n", + "bytes. Every checked-in asset is already a precomputed control, so the notebook sends it once as\n", + "base64 media inside JSON `extra_params`; no top-level reference upload is needed. TensorRT-LLM\n", + "decodes the media back to bytes at the HTTP boundary.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-setup", + "metadata": {}, + "outputs": [], + "source": [ + "from pathlib import Path\n", + "import os\n", + "\n", + "\n", + "def find_repo_root(start: Path) -> Path:\n", + " for path in [start, *start.parents]:\n", + " if (path / \"README.md\").exists() and (path / \"cookbooks\").exists():\n", + " return path\n", + " return start\n", + "\n", + "\n", + "COSMOS_ROOT = find_repo_root(Path.cwd().resolve())\n", + "TRANSFER_ROOT = COSMOS_ROOT / \"cookbooks\" / \"cosmos3\" / \"generator\" / \"transfer\"\n", + "OUTPUT_ROOT = Path(\n", + " os.environ.get(\"COSMOS3_TRTLLM_TRANSFER_OUTPUT_ROOT\", TRANSFER_ROOT / \"outputs\" / \"notebooks\" / \"trt_llm\")\n", + ").resolve()\n", + "TRTLLM_BASE_URL = os.environ.get(\"COSMOS3_TRTLLM_BASE_URL\", \"http://localhost:8000\").rstrip(\"/\")\n", + "TRTLLM_API_KEY = os.environ.get(\"COSMOS3_TRTLLM_API_KEY\", \"tensorrt_llm\")\n", + "OUTPUT_ROOT.mkdir(parents=True, exist_ok=True)\n", + "\n", + "print(\"COSMOS_ROOT:\", COSMOS_ROOT)\n", + "print(\"TRANSFER_ROOT:\", TRANSFER_ROOT)\n", + "print(\"OUTPUT_ROOT:\", OUTPUT_ROOT)\n", + "print(\"TRTLLM_BASE_URL:\", TRTLLM_BASE_URL)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-matrix", + "metadata": {}, + "source": [ + "## Request matrix\n", + "\n", + "| Control | Transport | Frames / fps | Text / control guidance |\n", + "| --- | --- | ---: | ---: |\n", + "| Edge | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Blur | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Depth | Base64 precomputed control | 121 / 30 | 3.0 / 1.5 |\n", + "| Segmentation | Base64 precomputed control | 121 / 30 | 3.0 / 2.0 |\n", + "| WSM | Base64 precomputed control | 100 / 10 | 1.0 / 3.0 |\n", + "\n", + "The blur control is 4:3 and explicitly selects `1104x832`; the other controls are 16:9 and\n", + "explicitly select `1280x720`. TensorRT-LLM can\n", + "also derive the nearest supported bucket from the first precomputed control when neither dimension\n", + "is set. To compute edge or blur on the server instead, upload a raw source as multipart\n", + "`video_reference` and set the corresponding hint to `true`.\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-helpers", + "metadata": {}, + "outputs": [], + "source": [ + "import base64\n", + "import json\n", + "import time\n", + "\n", + "import requests\n", + "from IPython.display import Video, display\n", + "\n", + "\n", + "def server_root_url() -> str:\n", + " return TRTLLM_BASE_URL[:-3] if TRTLLM_BASE_URL.endswith(\"/v1\") else TRTLLM_BASE_URL\n", + "\n", + "\n", + "def video_api_url() -> str:\n", + " root = TRTLLM_BASE_URL if TRTLLM_BASE_URL.endswith(\"/v1\") else f\"{TRTLLM_BASE_URL}/v1\"\n", + " return f\"{root}/videos/sync\"\n", + "\n", + "\n", + "def wait_for_server(timeout_s: int = 1800, interval_s: int = 10) -> None:\n", + " deadline = time.time() + timeout_s\n", + " url = f\"{server_root_url()}/health\"\n", + " while time.time() < deadline:\n", + " try:\n", + " response = requests.get(url, timeout=10)\n", + " if response.ok:\n", + " print(\"TensorRT-LLM server is ready:\", url)\n", + " return\n", + " except requests.RequestException as exc:\n", + " print(\"waiting for TensorRT-LLM:\", exc)\n", + " time.sleep(interval_s)\n", + " raise TimeoutError(f\"TensorRT-LLM did not become ready at {url}\")\n", + "\n", + "\n", + "CONTROL_CASES = {\n", + " \"edge\": {\n", + " \"control\": \"assets/edge/control_edge.mp4\",\n", + " \"prompt\": \"assets/edge/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " \"size\": \"1280x720\",\n", + " \"hint\": {\"preset_edge_threshold\": \"medium\"},\n", + " },\n", + " \"blur\": {\n", + " \"control\": \"assets/blur/control_blur.mp4\",\n", + " \"prompt\": \"assets/blur/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " \"size\": \"1104x832\",\n", + " \"hint\": {\"preset_blur_strength\": \"medium\"},\n", + " },\n", + " \"depth\": {\n", + " \"control\": \"assets/depth/control_depth.mp4\",\n", + " \"prompt\": \"assets/depth/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 1.5,\n", + " \"size\": \"1280x720\",\n", + " },\n", + " \"seg\": {\n", + " \"control\": \"assets/seg/control_seg.mp4\",\n", + " \"prompt\": \"assets/seg/prompt.json\",\n", + " \"num_frames\": 121,\n", + " \"fps\": 30,\n", + " \"guidance_scale\": 3.0,\n", + " \"control_guidance\": 2.0,\n", + " \"size\": \"1280x720\",\n", + " },\n", + " \"wsm\": {\n", + " \"control\": \"assets/wsm/control_wsm.mp4\",\n", + " \"prompt\": \"assets/wsm/prompt.json\",\n", + " \"num_frames\": 100,\n", + " \"fps\": 10,\n", + " \"guidance_scale\": 1.0,\n", + " \"control_guidance\": 3.0,\n", + " \"size\": \"1280x720\",\n", + " },\n", + "}\n", + "\n", + "\n", + "def compact_json(path: Path) -> str:\n", + " return json.dumps(json.loads(path.read_text()), ensure_ascii=True, separators=(\",\", \":\"))\n", + "\n", + "\n", + "def run_transfer(control_name: str) -> Path:\n", + " \"\"\"Send one precomputed control through the JSON/base64 input contract.\"\"\"\n", + " spec = CONTROL_CASES[control_name]\n", + " control_path = (TRANSFER_ROOT / spec[\"control\"]).resolve()\n", + " prompt_path = (TRANSFER_ROOT / spec[\"prompt\"]).resolve()\n", + " negative_path = TRANSFER_ROOT / \"assets\" / \"negative_prompt.json\"\n", + " for path in (control_path, prompt_path, negative_path):\n", + " if not path.exists():\n", + " raise FileNotFoundError(path)\n", + "\n", + " hint = dict(spec.get(\"hint\", {}))\n", + " # These files are already edge/blur/depth/seg/WSM control representations.\n", + " # Send each one directly instead of also treating it as a raw input video.\n", + " hint[\"control\"] = base64.b64encode(control_path.read_bytes()).decode(\"ascii\")\n", + "\n", + " extra_params = {\n", + " \"use_duration_template\": False,\n", + " \"use_resolution_template\": False,\n", + " \"use_system_prompt\": False,\n", + " \"use_guardrails\": True,\n", + " \"flow_shift\": 10.0,\n", + " control_name: hint,\n", + " \"control_guidance\": spec[\"control_guidance\"],\n", + " \"num_video_frames_per_chunk\": spec[\"num_frames\"],\n", + " \"num_conditional_frames\": 1,\n", + " \"num_first_chunk_conditional_frames\": 0,\n", + " \"share_vision_temporal_positions\": True,\n", + " \"emphasize_control_in_prompt\": False,\n", + " \"max_frames\": spec[\"num_frames\"],\n", + " }\n", + " request_body = {\n", + " \"prompt\": compact_json(prompt_path),\n", + " \"negative_prompt\": compact_json(negative_path),\n", + " \"size\": spec[\"size\"],\n", + " \"num_frames\": spec[\"num_frames\"],\n", + " \"fps\": spec[\"fps\"],\n", + " \"num_inference_steps\": 50,\n", + " \"guidance_scale\": spec[\"guidance_scale\"],\n", + " \"max_sequence_length\": 4096,\n", + " \"seed\": 2026,\n", + " \"format\": \"mp4\",\n", + " \"response_format\": \"file\",\n", + " \"extra_params\": extra_params,\n", + " }\n", + " headers = {\"Accept\": \"video/mp4\"}\n", + " if TRTLLM_API_KEY:\n", + " headers[\"Authorization\"] = f\"Bearer {TRTLLM_API_KEY}\"\n", + " response = requests.post(\n", + " video_api_url(),\n", + " json=request_body,\n", + " headers=headers,\n", + " timeout=3600,\n", + " )\n", + " if not response.ok:\n", + " raise RuntimeError(f\"TensorRT-LLM request failed ({response.status_code}): {response.text}\")\n", + " content_type = response.headers.get(\"content-type\", \"\").lower()\n", + " if not response.content:\n", + " raise RuntimeError(\"TensorRT-LLM returned an empty media response\")\n", + " if \"video/mp4\" not in content_type or response.content[4:8] != b\"ftyp\":\n", + " raise RuntimeError(f\"Expected browser-compatible MP4, got {content_type!r}\")\n", + " output_path = OUTPUT_ROOT / f\"transfer_{control_name}.mp4\"\n", + " output_path.write_bytes(response.content)\n", + " print(\"control:\", control_path)\n", + " print(\"extra_params:\", json.dumps({**extra_params, control_name: \"\"}, indent=2))\n", + " print(\"saved:\", output_path, content_type)\n", + " return output_path\n", + "\n", + "\n", + "def view_transfer(path: Path) -> None:\n", + " display(Video(str(path), embed=True))\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-edge-title", + "metadata": {}, + "source": [ + "## Canny / Edge transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-edge-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "edge_output = run_transfer(\"edge\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-edge-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(edge_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-blur-title", + "metadata": {}, + "source": [ + "## Blur transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-blur-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "blur_output = run_transfer(\"blur\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-blur-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(blur_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-depth-title", + "metadata": {}, + "source": [ + "## Depth transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-depth-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "depth_output = run_transfer(\"depth\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-depth-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(depth_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-seg-title", + "metadata": {}, + "source": [ + "## Segmentation transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-seg-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "seg_output = run_transfer(\"seg\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-seg-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(seg_output)\n" + ] + }, + { + "cell_type": "markdown", + "id": "transfer-trt-wsm-title", + "metadata": {}, + "source": [ + "## World Scenario Model (WSM) transfer\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-wsm-run", + "metadata": {}, + "outputs": [], + "source": [ + "wait_for_server()\n", + "wsm_output = run_transfer(\"wsm\")\n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "transfer-trt-wsm-view", + "metadata": {}, + "outputs": [], + "source": [ + "view_transfer(wsm_output)\n" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/cookbooks/cosmos3/generator/transfer/specs/wsm.json b/cookbooks/cosmos3/generator/transfer/specs/wsm.json index eb18d535..e56a7b60 100644 --- a/cookbooks/cosmos3/generator/transfer/specs/wsm.json +++ b/cookbooks/cosmos3/generator/transfer/specs/wsm.json @@ -3,12 +3,12 @@ "model_mode": "video2video", "resolution": "720", "aspect_ratio": "16,9", - "num_frames": 101, + "num_frames": 100, "fps": 10, "shift": 10.0, "num_steps": 50, "seed": 2026, - "num_video_frames_per_chunk": 101, + "num_video_frames_per_chunk": 100, "num_conditional_frames": 1, "num_first_chunk_conditional_frames": 0, "share_vision_temporal_positions": true,