feat(vllm): install the PD router into the released CUDA 0.27.1 images - #277
Conversation
The router landed in pack/cuda/Dockerfile.vllm (5c49beb) after 0.27.1 was built, so a disaggregated group deployed on those images has no router to run; 0.29.0 ships it already. CANN needs no operation -- 0.23.0 is still what the matrix builds and its recipe carries the router, so a normal release republishes it.
There was a problem hiding this comment.
Code Review
This pull request introduces a post-operation to install the vllm-router package for vLLM 0.27.1 CUDA images, addressing a missing component in previously released images. The review feedback highlights that the DEPENDENCY_PACKAGES argument is currently empty and should be populated with the new dependencies (vllm-router, fastapi, uvicorn, aiohttp, orjson, requests, and setproctitle) to ensure the runner.py.json dependency file remains accurate.
|
|
||
| ## Probe Dependencies | ||
|
|
||
| ARG DEPENDENCY_PACKAGES="" |
There was a problem hiding this comment.
According to pack/.post_operation/README.md, new operations that modify the Python environment must probe dependencies to keep runner.py.json up to date. Since this operation installs vllm-router and its dependencies, DEPENDENCY_PACKAGES should be populated with these new packages instead of being empty.
The dependencies for vllm-router are mentioned in the comments as fastapi, uvicorn, aiohttp, orjson, requests, and setproctitle. These, along with vllm-router itself, should be added to DEPENDENCY_PACKAGES to ensure they are correctly recorded.
ARG DEPENDENCY_PACKAGES="vllm-router fastapi uvicorn aiohttp orjson requests setproctitle"
The router landed in pack/cuda/Dockerfile.vllm (5c49beb) after 0.27.1 was built, so a disaggregated group deployed on those images has no router to run; 0.29.0 ships it already.