Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 61 additions & 0 deletions pack/.post_operation/20260920_vllm_install_router/cuda/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
ARG CMAKE_MAX_JOBS
ARG CUDA_VERSION=12.9
ARG VLLM_VERSION=0.27.1
ARG VLLM_ROUTER_VERSION=0.1.15

FROM gpustack/runner:cuda${CUDA_VERSION}-vllm${VLLM_VERSION} AS vllm
SHELL ["/bin/bash", "-eo", "pipefail", "-c"]

ARG TARGETPLATFORM
ARG TARGETOS
ARG TARGETARCH

## Install vLLM Router
##
## `pack/cuda/Dockerfile.vllm` only gained the router after 0.27.1 was cut, so the released 0.27.1
## images carry none of it: a disaggregated group deployed on them has no router to run, and either
## falls back to an engine example script (no metrics, no circuit breaker, no health check) or does
## not start at all. 0.29.0 already ships it, which is why 0.27.1 is the only CUDA runner here.
##
## Pinned to the same 0.1.15 the recipe installs, so that a 0.27.1 image and a 0.29.0 image render
## the same router. The wheel is prebuilt for both aarch64 and x86_64, and it depends on nothing
## but pure-Python web packages (fastapi, uvicorn, aiohttp, orjson, requests, setproctitle) -- it
## does not depend on vLLM, so this install cannot move the torch/vLLM stack already in the image.

ARG VLLM_ROUTER_VERSION

RUN <<EOF
# vLLM Router

uv pip install \
vllm-router==${VLLM_ROUTER_VERSION}

# Fail the build rather than ship an image whose router is missing: the
# alternative surfaces as a container that exits at deploy time, one
# layer away from anything that explains it.
command -v vllm-router >/dev/null

# Review
uv pip tree

# Cleanup
rm -rf /var/tmp/* \
&& rm -rf /tmp/*
EOF

## Probe Dependencies

ARG DEPENDENCY_PACKAGES=""

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

According to pack/.post_operation/README.md, new operations that modify the Python environment must probe dependencies to keep runner.py.json up to date. Since this operation installs vllm-router and its dependencies, DEPENDENCY_PACKAGES should be populated with these new packages instead of being empty.

The dependencies for vllm-router are mentioned in the comments as fastapi, uvicorn, aiohttp, orjson, requests, and setproctitle. These, along with vllm-router itself, should be added to DEPENDENCY_PACKAGES to ensure they are correctly recorded.

ARG DEPENDENCY_PACKAGES="vllm-router fastapi uvicorn aiohttp orjson requests setproctitle"

RUN --mount=type=bind,from=shared,source=probe_dependencies.sh,target=/tmp/probe_dependencies.sh \
DEPENDENCY_PACKAGES="${DEPENDENCY_PACKAGES}" bash /tmp/probe_dependencies.sh

## Entrypoint

WORKDIR /
ENTRYPOINT [ "tini", "--" ]

## Export Dependencies

FROM scratch AS vllm-deps

COPY --from=vllm /etc/gpustack-runner/dependencies.json /
32 changes: 32 additions & 0 deletions pack/.post_operation/20260920_vllm_install_router/matrix.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
rules:

# CUDA 0.27.1 only, and deliberately so.
#
# The router install lives in `pack/cuda/Dockerfile.vllm` and `pack/cann/Dockerfile.vllm` -- the
# two recipes GPUStack renders a PD mode for. Of the released images that miss it:
#
# - cuda 0.27.1 is no longer in `pack/matrix.yaml` (CUDA stands at 0.29.0, which already ships
# the router), so rewriting the tag here is the only way to put a router into it.
# - cann 0.23.0 is still exactly what `pack/matrix.yaml` builds, and the recipe already carries
# the router, so a normal `for_release` pack of backend `cann` rebuilds those four variants
# with it. That is a plain release rather than a mutation, so it does not belong here.
# - rocm never installed a router at all; its 0.29.0 image has none either, and patching only
# 0.27.1 would leave the newer image the odd one out.

## Packed NVIDIA CUDA 13.0.
##
- backend: "cuda"
services:
- "vllm"
args:
- "CUDA_VERSION=13.0"
- "VLLM_VERSION=0.27.1"

## Packed NVIDIA CUDA 12.9.
##
- backend: "cuda"
services:
- "vllm"
args:
- "CUDA_VERSION=12.9"
- "VLLM_VERSION=0.27.1"
1 change: 1 addition & 0 deletions pack/.post_operation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,3 +101,4 @@ mutated stay as they were until the next release rebuilds those images.
- [x] 2026-03-03: Fix malformed ARM64 image for vLLM 0.15.1 of CUDA released images.
- [x] 2026-09-01: Pin `numpy` to 1.26.4 and remove CUDA-only NIXL EP packages for vLLM 0.18.1 of DTK 26.04 released images.
- [x] 2026-09-16: Patch vLLM 0.24.0/0.25.1/0.27.1/0.29.0 of CUDA/ROCm released images to fix mooncake prom metrics issue.
- [ ] 2026-09-20: Install `vllm-router` package for vLLM 0.27.1 of CUDA released images.
Loading