-
Notifications
You must be signed in to change notification settings - Fork 16
feat(vllm): install the PD router into the released CUDA 0.27.1 images #277
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
61 changes: 61 additions & 0 deletions
61
pack/.post_operation/20260920_vllm_install_router/cuda/Dockerfile
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| ARG CMAKE_MAX_JOBS | ||
| ARG CUDA_VERSION=12.9 | ||
| ARG VLLM_VERSION=0.27.1 | ||
| ARG VLLM_ROUTER_VERSION=0.1.15 | ||
|
|
||
| FROM gpustack/runner:cuda${CUDA_VERSION}-vllm${VLLM_VERSION} AS vllm | ||
| SHELL ["/bin/bash", "-eo", "pipefail", "-c"] | ||
|
|
||
| ARG TARGETPLATFORM | ||
| ARG TARGETOS | ||
| ARG TARGETARCH | ||
|
|
||
| ## Install vLLM Router | ||
| ## | ||
| ## `pack/cuda/Dockerfile.vllm` only gained the router after 0.27.1 was cut, so the released 0.27.1 | ||
| ## images carry none of it: a disaggregated group deployed on them has no router to run, and either | ||
| ## falls back to an engine example script (no metrics, no circuit breaker, no health check) or does | ||
| ## not start at all. 0.29.0 already ships it, which is why 0.27.1 is the only CUDA runner here. | ||
| ## | ||
| ## Pinned to the same 0.1.15 the recipe installs, so that a 0.27.1 image and a 0.29.0 image render | ||
| ## the same router. The wheel is prebuilt for both aarch64 and x86_64, and it depends on nothing | ||
| ## but pure-Python web packages (fastapi, uvicorn, aiohttp, orjson, requests, setproctitle) -- it | ||
| ## does not depend on vLLM, so this install cannot move the torch/vLLM stack already in the image. | ||
|
|
||
| ARG VLLM_ROUTER_VERSION | ||
|
|
||
| RUN <<EOF | ||
| # vLLM Router | ||
|
|
||
| uv pip install \ | ||
| vllm-router==${VLLM_ROUTER_VERSION} | ||
|
|
||
| # Fail the build rather than ship an image whose router is missing: the | ||
| # alternative surfaces as a container that exits at deploy time, one | ||
| # layer away from anything that explains it. | ||
| command -v vllm-router >/dev/null | ||
|
|
||
| # Review | ||
| uv pip tree | ||
|
|
||
| # Cleanup | ||
| rm -rf /var/tmp/* \ | ||
| && rm -rf /tmp/* | ||
| EOF | ||
|
|
||
| ## Probe Dependencies | ||
|
|
||
| ARG DEPENDENCY_PACKAGES="" | ||
| RUN --mount=type=bind,from=shared,source=probe_dependencies.sh,target=/tmp/probe_dependencies.sh \ | ||
| DEPENDENCY_PACKAGES="${DEPENDENCY_PACKAGES}" bash /tmp/probe_dependencies.sh | ||
|
|
||
| ## Entrypoint | ||
|
|
||
| WORKDIR / | ||
| ENTRYPOINT [ "tini", "--" ] | ||
|
|
||
| ## Export Dependencies | ||
|
|
||
| FROM scratch AS vllm-deps | ||
|
|
||
| COPY --from=vllm /etc/gpustack-runner/dependencies.json / | ||
32 changes: 32 additions & 0 deletions
32
pack/.post_operation/20260920_vllm_install_router/matrix.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,32 @@ | ||
| rules: | ||
|
|
||
| # CUDA 0.27.1 only, and deliberately so. | ||
| # | ||
| # The router install lives in `pack/cuda/Dockerfile.vllm` and `pack/cann/Dockerfile.vllm` -- the | ||
| # two recipes GPUStack renders a PD mode for. Of the released images that miss it: | ||
| # | ||
| # - cuda 0.27.1 is no longer in `pack/matrix.yaml` (CUDA stands at 0.29.0, which already ships | ||
| # the router), so rewriting the tag here is the only way to put a router into it. | ||
| # - cann 0.23.0 is still exactly what `pack/matrix.yaml` builds, and the recipe already carries | ||
| # the router, so a normal `for_release` pack of backend `cann` rebuilds those four variants | ||
| # with it. That is a plain release rather than a mutation, so it does not belong here. | ||
| # - rocm never installed a router at all; its 0.29.0 image has none either, and patching only | ||
| # 0.27.1 would leave the newer image the odd one out. | ||
|
|
||
| ## Packed NVIDIA CUDA 13.0. | ||
| ## | ||
| - backend: "cuda" | ||
| services: | ||
| - "vllm" | ||
| args: | ||
| - "CUDA_VERSION=13.0" | ||
| - "VLLM_VERSION=0.27.1" | ||
|
|
||
| ## Packed NVIDIA CUDA 12.9. | ||
| ## | ||
| - backend: "cuda" | ||
| services: | ||
| - "vllm" | ||
| args: | ||
| - "CUDA_VERSION=12.9" | ||
| - "VLLM_VERSION=0.27.1" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
According to
pack/.post_operation/README.md, new operations that modify the Python environment must probe dependencies to keeprunner.py.jsonup to date. Since this operation installsvllm-routerand its dependencies,DEPENDENCY_PACKAGESshould be populated with these new packages instead of being empty.The dependencies for
vllm-routerare mentioned in the comments asfastapi,uvicorn,aiohttp,orjson,requests, andsetproctitle. These, along withvllm-routeritself, should be added toDEPENDENCY_PACKAGESto ensure they are correctly recorded.