Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,8 +56,8 @@ The following table lists the supported accelerated backends and their correspon

| CUDA Version <br/> (Variant) | vLLM | SGLang | VoxBox |
|------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------|----------|
| 13.0 | **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1`,<br/> `0.18.1` | `0.5.18`, `0.5.15.post1`, <br/>`0.5.14`, `0.5.12.post1` | |
| 12.9 | **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1`,<br/> `0.18.1`, `0.17.1`,<br/> `0.16.0`, `0.15.1`,<br/> `0.14.1`, `0.13.0`,<br/> `0.12.0`, `0.11.2` | `0.5.18`, `0.5.15.post1`,<br/> `0.5.14`, `0.5.12.post1`, <br/>`0.5.9`, `0.5.8.post1`, <br/>`0.5.7`, `0.5.6.post2` | |
| 13.0 | **`0.30.0`**, **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1`,<br/> `0.18.1` | `0.5.18`, `0.5.15.post1`, <br/>`0.5.14`, `0.5.12.post1` | |
| 12.9 | **`0.30.0`**, **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1`,<br/> `0.18.1`, `0.17.1`,<br/> `0.16.0`, `0.15.1`,<br/> `0.14.1`, `0.13.0`,<br/> `0.12.0`, `0.11.2` | `0.5.18`, `0.5.15.post1`,<br/> `0.5.14`, `0.5.12.post1`, <br/>`0.5.9`, `0.5.8.post1`, <br/>`0.5.7`, `0.5.6.post2` | |
| 12.8 | `0.17.1`, `0.16.0`, <br/>`0.15.1`, `0.14.1`, <br/>`0.13.0`, `0.12.0`, <br/>`0.11.2`, `0.10.2` | `0.5.9`, `0.5.8.post1`, <br/>`0.5.7`, `0.5.6.post2`, <br/>`0.5.5.post3` | `0.0.21` |
| 12.6 | `0.15.1`, `0.14.1`, <br/>`0.13.0`, `0.12.0`, <br/>`0.11.2`, `0.10.2` | | `0.0.21` |

Expand Down Expand Up @@ -96,7 +96,7 @@ The following table lists the supported accelerated backends and their correspon

| ROCm Version <br/> (Variant) | vLLM | SGLang |
|------------------------------|---------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------|
| 7.2 | **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1` | `0.5.18`, `0.5.15.post1`, <br/>`0.5.14`, `0.5.12.post1` |
| 7.2 | **`0.30.0`**, **`0.29.0`**, **`0.27.1`**,<br/> **`0.25.1`**, **`0.24.0`**,<br/> `0.22.1`, `0.21.0`,<br/> `0.20.2`, `0.19.1` | `0.5.18`, `0.5.15.post1`, <br/>`0.5.14`, `0.5.12.post1` |
| 7.1 | `0.17.1` | |
| 7.0 | `0.18.1`, `0.16.0`,<br/> `0.15.1`, `0.14.1`,<br/> `0.13.0`, `0.12.0`,<br/> `0.11.2` | `0.5.9`, `0.5.8.post1`, <br/>`0.5.7`, `0.5.6.post2` |
| 6.4 | `0.16.0`, `0.15.1`,<br/> `0.14.1`, `0.13.0`,<br/> `0.12.0`, `0.11.2`,<br/> `0.10.2` | `0.5.8.post1`, `0.5.7`, <br/>`0.5.6.post2`, `0.5.5.post3` |
Expand Down
8 changes: 6 additions & 2 deletions pack/cuda/Dockerfile.sglang
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ ARG SGLANG_TORCH_CUDA_VERSION=${CUDA_VERSION}
## Keep >= 0.4.6 and in sync with the ROCm and vLLM pins: SGLang's lmc_radix_cache.py imports
## `lmcache.integration.sglang.multi_process_adapter`, added in 0.4.6, and the MP wire protocol
## has no version handshake. The SGLang base ships no lmcache, so it is always built from source.
ARG SGLANG_LMCACHE_VERSION=0.5.4
ARG SGLANG_LMCACHE_VERSION=0.5.5
ARG SGLANG_NVIDIA_HPCX_VERSION=2.24.1_cuda13
ARG SGLANG_AWS_EFA_VERSION=1.46.0

Expand Down Expand Up @@ -192,6 +192,10 @@ RUN <<EOF
## Note: don't pin cuda-bindings —— cu129 base has 12.9.4, cu130 base has 13.2.0,
## both compatible with their respective torch. Hardcoding 12.9.4 on cu130 downgrades
## and breaks `torch -> cuda-bindings<14,>=13.0.3` requirement.
## 0.5.5 generates its gRPC stubs at build time (setup.py _BuildPyWithGrpcStubs);
## provide the pinned codegen tools its build-system requirements name, since
## `--no-isolation` skips installing them.
uv pip install grpcio==1.84.0 grpcio-tools==1.84.0
pushd /tmp/lmcache \
&& sed -i "s/\"torch==.*\"/\"torch\"/g" /tmp/lmcache/pyproject.toml \
&& cat /tmp/lmcache/pyproject.toml \
Expand Down Expand Up @@ -373,7 +377,7 @@ RUN <<EOF
# both share the same cupy/ tree, so uninstalling one breaks the other. Purge both, then
# reinstall only the variant matching the CUDA major we build for, at the version already
# present in the image. Since 0.4.6 lmcache splits its requirements per CUDA major, so our
# 0.5.4 pin agrees with this reinstall (see LMCACHE_CUDA_MAJOR in the build stage above),
# 0.5.5 pin agrees with this reinstall (see LMCACHE_CUDA_MAJOR in the build stage above),
# leaving MSCCL++ as the only source that has to be corrected here.

IFS="." read -r CUDA_MAJOR CUDA_MINOR CUDA_PATCH <<< "${CUDA_VERSION}"
Expand Down
71 changes: 65 additions & 6 deletions pack/cuda/Dockerfile.vllm
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,22 @@ ARG PYTHON_VERSION=3.12
ARG CMAKE_MAX_JOBS
ARG CUDA_VERSION=13.0.1
ARG CUDA_ARCHS
ARG VLLM_BASE_IMAGE=vllm/vllm-openai:v0.29.0-ubuntu2404
ARG VLLM_VERSION=0.29.0
ARG VLLM_TORCH_VERSION=2.13.0
ARG VLLM_BASE_IMAGE=vllm/vllm-openai:v0.30.0-ubuntu2404
ARG VLLM_VERSION=0.30.0
ARG VLLM_TORCH_VERSION=2.14.0
## torchvision released in lockstep with torch (0.29.0 pairs with 2.14.0); torchaudio
## has no 2.12.0 release, so the base's 2.11.0 stays on both cu129 and cu130.
ARG VLLM_TORCHVISION_VERSION=0.29.0
ARG VLLM_TORCH_CUDA_VERSION=${CUDA_VERSION}
## Built from source, replacing the base's own lmcache: that one is a PyPI wheel compiled
## against lmcache's build-time torch pin, which need not match the torch this base ships, and
## on a mismatch lmcache silently degrades to its torch baseline (vllm-project/vllm#53424).
## Keep in sync with the ROCm and SGLang pins —— the MP wire protocol has no version handshake.
ARG VLLM_LMCACHE_VERSION=0.5.4
ARG VLLM_LMCACHE_VERSION=0.5.5
ARG VLLM_ROUTER_VERSION=0.1.15
ARG VLLM_OMNI_COMMIT=aff7d64948f6
## A tag, not a commit: v0.30.0rc1 is the omni release rebased onto vLLM 0.30.0
## (the previous commit pinned a 0.29.0-era tree).
ARG VLLM_OMNI_COMMIT=v0.30.0rc1

FROM ${VLLM_BASE_IMAGE} AS vllm-build
SHELL ["/bin/bash", "-eo", "pipefail", "-c"]
Expand Down Expand Up @@ -121,12 +126,56 @@ ENV UV_SYSTEM_PYTHON=1 \

ARG VLLM_VERSION
ARG VLLM_TORCH_VERSION
ARG VLLM_TORCHVISION_VERSION
ARG VLLM_TORCH_CUDA_VERSION

ENV VLLM_VERSION=${VLLM_VERSION} \
VLLM_TORCH_VERSION=${VLLM_TORCH_VERSION} \
VLLM_TORCH_CUDA_VERSION=${VLLM_TORCH_CUDA_VERSION}

## Upgrade Torch
#
# The base image ships the torch it was built against (2.13.0 for the v0.30.0
# bases). When a rule targets a newer torch, install it from the matching
# PyTorch wheel index; the base's prebuilt wheels (vllm, flashinfer, fa3-fwd)
# were compiled against the older torch and rely on torch's C++ ABI staying
# stable within the major version. A rule whose base already ships the target
# torch (cu129: 2.14.0 is published for cu130 only) passes VLLM_TORCH_VERSION
# equal to the base's, and the step is a no-op.

ARG VLLM_TORCH_VERSION
ARG VLLM_TORCHVISION_VERSION
ARG VLLM_TORCH_CUDA_VERSION

RUN <<EOF
# Torch

INSTALLED_TORCH="$(python -c 'import torch; print(torch.__version__)')"
if [[ "${INSTALLED_TORCH}" == "${VLLM_TORCH_VERSION}"* ]]; then
echo "Torch ${INSTALLED_TORCH} already matches ${VLLM_TORCH_VERSION}; skipping the upgrade."
exit 0
fi

IFS="." read -r CUDA_MAJOR CUDA_MINOR CUDA_PATCH <<< "${VLLM_TORCH_CUDA_VERSION}"

## `--extra-index-url`, not `--index-url`: the PyTorch index carries torch/torchvision
## only, while the transitive deps (filelock, sympy, ...) and the nvidia-* wheels
## resolve from PyPI. The `+cu${CUDA_MAJOR}${CUDA_MINOR}` local versions pin the
## index wheels exactly, so the PyPI copy of the same version can never win.
uv pip install \
"torch==${VLLM_TORCH_VERSION}+cu${CUDA_MAJOR}${CUDA_MINOR}" \
"torchvision==${VLLM_TORCHVISION_VERSION}+cu${CUDA_MAJOR}${CUDA_MINOR}" \
--extra-index-url "https://download.pytorch.org/whl/cu${CUDA_MAJOR}${CUDA_MINOR}"

# Review
uv pip show torch torchvision
python -c 'import torch; print("torch", torch.__version__, "cuda", torch.version.cuda)'

# Cleanup
rm -rf /var/tmp/* \
&& rm -rf /tmp/*
EOF


FROM vllm-build AS vllm-build-omni
SHELL ["/bin/bash", "-eo", "pipefail", "-c"]
Expand Down Expand Up @@ -260,10 +309,20 @@ RUN <<EOF
## and breaks `torch -> cuda-bindings<14,>=13.0.3` requirement.
## Note: separate statements, not an `&&` chain —— `bash -e` ignores a failing non-final
## command of a list, which would leave /workspace empty and fail later at install time.
## 0.5.5 generates its gRPC stubs at build time (setup.py _BuildPyWithGrpcStubs);
## provide the pinned codegen tools its build-system requirements name, since
## `--no-isolation` skips installing them.
uv pip install grpcio==1.84.0 grpcio-tools==1.84.0
cd /tmp/lmcache
## Drop the build-time torch pin: we compile against the base's own torch on purpose.
sed -i "s/\"torch==.*\"/\"torch\"/g" /tmp/lmcache/pyproject.toml
cat /tmp/lmcache/pyproject.toml
## torch 2.14's headers require C++20 (#error in torch/all.h), but lmcache 0.5.5
## hardcodes -std=c++17 for its g++ sources (pybind.cpp); nvcc already defaults
## to C++20 under CUDA 13. Harmless on the cu129 base's torch 2.13.
sed -i 's/-std=c++17/-std=c++20/g' \
/tmp/lmcache/setup_extensions/common_cpp.py \
/tmp/lmcache/setup_extensions/build_profiles/cuda.py
## `--skip-dependency-check`: the base does not ship torch's full dependency closure.
python -m build --no-isolation --skip-dependency-check --wheel
tree -hs /tmp/lmcache/dist
Expand Down Expand Up @@ -473,7 +532,7 @@ soundfile
mistral_common[audio]

# tokenizer extras
fastokens==0.2.0
fastokens==0.3.2
EOT
uv pip install \
-r /tmp/requirements.txt
Expand Down
22 changes: 14 additions & 8 deletions pack/matrix.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -101,8 +101,8 @@ rules:
- "linux/amd64"
args:
- "ROCM_VERSION=7.2.3"
- "VLLM_VERSION=0.29.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0"
- "VLLM_VERSION=0.30.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai-rocm:v0.30.0"
## AMD ROCm 7.2.1 - SGLang 0.5.18 (rocm720-mi30x)
##
- backend: "rocm"
Expand All @@ -119,15 +119,17 @@ rules:
# NVIDIA CUDA
#

## NVIDIA CUDA 13.0.1
## NVIDIA CUDA 13.0.3
## (torch is upgraded from the base's 2.13.0 to 2.14.0 —— see the
## "Upgrade Torch" step in pack/cuda/Dockerfile.vllm)
##
- backend: "cuda"
services:
- "vllm"
args:
- "CUDA_VERSION=13.0.1"
- "VLLM_VERSION=0.29.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai:v0.29.0-ubuntu2404"
- "CUDA_VERSION=13.0.3"
- "VLLM_VERSION=0.30.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai:v0.30.0-ubuntu2404"
## NVIDIA CUDA 13.0.1 - SGLang 0.5.18
##
- backend: "cuda"
Expand All @@ -138,14 +140,18 @@ rules:
- "SGLANG_VERSION=0.5.18"
- "SGLANG_BASE_IMAGE=lmsysorg/sglang:v0.5.18-cu130"
## NVIDIA CUDA 12.9.1
## (stays on the base's torch 2.13.0: 2.14.0 is published for cu130 only,
## so VLLM_TORCH_VERSION is pinned to the base's version and the
## "Upgrade Torch" step in pack/cuda/Dockerfile.vllm is a no-op)
##
- backend: "cuda"
services:
- "vllm"
args:
- "CUDA_VERSION=12.9.1"
- "VLLM_VERSION=0.29.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai:v0.29.0-cu129-ubuntu2404"
- "VLLM_VERSION=0.30.0"
- "VLLM_BASE_IMAGE=vllm/vllm-openai:v0.30.0-cu129-ubuntu2404"
- "VLLM_TORCH_VERSION=2.13.0"
## NVIDIA CUDA 12.9.1 - SGLang 0.5.18
##
- backend: "cuda"
Expand Down
6 changes: 5 additions & 1 deletion pack/rocm/Dockerfile.sglang
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ ARG SGLANG_TORCH_VERSION=2.9.1
ARG SGLANG_TORCH_ROCM_VERSION=${ROCM_VERSION}
## Keep >= 0.4.6 and in sync with the CUDA pin —— see pack/cuda/Dockerfile.sglang.
## Always built from source: upstream's ROCm wheel targets torch 2.11.0+rocm7.2.
ARG SGLANG_LMCACHE_VERSION=0.5.4
ARG SGLANG_LMCACHE_VERSION=0.5.5

FROM ${SGLANG_BASE_IMAGE} AS sglang-build
SHELL ["/bin/bash", "-eo", "pipefail", "-c"]
Expand Down Expand Up @@ -198,6 +198,10 @@ RUN <<EOF
git -C /tmp clone --recursive --shallow-submodules \
--depth 1 --branch v${SGLANG_LMCACHE_VERSION} --single-branch \
https://github.com/LMCache/LMCache.git lmcache
## 0.5.5 generates its gRPC stubs at build time (setup.py _BuildPyWithGrpcStubs);
## provide the pinned codegen tools its build-system requirements name, since
## `--no-isolation` skips installing them.
uv pip install grpcio==1.84.0 grpcio-tools==1.84.0
pushd /tmp/lmcache \
&& sed -i "s/\"torch==.*\"/\"torch\"/g" /tmp/lmcache/pyproject.toml \
&& cat /tmp/lmcache/pyproject.toml \
Expand Down
18 changes: 13 additions & 5 deletions pack/rocm/Dockerfile.vllm
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,21 @@ ARG PYTHON_VERSION=3.12
ARG CMAKE_MAX_JOBS
ARG ROCM_VERSION=7.2.3
ARG ROCM_ARCHS
ARG VLLM_BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0
ARG VLLM_VERSION=0.29.0
ARG VLLM_BASE_IMAGE=vllm/vllm-openai-rocm:v0.30.0
ARG VLLM_VERSION=0.30.0
## The v0.30.0 ROCm base is built from the ROCm pytorch release/2.12 branch,
## so it keeps the base's torch; there is no ROCm equivalent of the CUDA
## "Upgrade Torch" step.
ARG VLLM_TORCH_VERSION=2.12.0
ARG VLLM_TORCH_ROCM_VERSION=${ROCM_VERSION}
## The ROCm base ships no lmcache of its own, so it is always built from source here.
## Keep in sync with the CUDA and SGLang pins —— see pack/cuda/Dockerfile.vllm.
ARG VLLM_LMCACHE_VERSION=0.5.4
ARG VLLM_LMCACHE_VERSION=0.5.5
## Upstream's own HIP build —— same source and CMake options we used to compile here,
## but produced in a clean rocm/dev-ubuntu-22.04 and smoke-tested before release.
ARG VLLM_MOONCAKE_VERSION=0.3.13.post1
ARG VLLM_OMNI_COMMIT=aff7d64948f6
## A tag, not a commit: v0.30.0rc1 is the omni release rebased onto vLLM 0.30.0.
ARG VLLM_OMNI_COMMIT=v0.30.0rc1

FROM ${VLLM_BASE_IMAGE} AS vllm-build
SHELL ["/bin/bash", "-eo", "pipefail", "-c"]
Expand Down Expand Up @@ -185,6 +189,10 @@ RUN <<EOF
git -C /tmp clone --recursive --shallow-submodules \
--depth 1 --branch v${VLLM_LMCACHE_VERSION} --single-branch \
https://github.com/LMCache/LMCache.git lmcache
## 0.5.5 generates its gRPC stubs at build time (setup.py _BuildPyWithGrpcStubs);
## provide the pinned codegen tools its build-system requirements name, since
## `--no-isolation` skips installing them.
uv pip install grpcio==1.84.0 grpcio-tools==1.84.0
pushd /tmp/lmcache \
&& sed -i "s/\"torch==.*\"/\"torch\"/g" /tmp/lmcache/pyproject.toml \
&& cat /tmp/lmcache/pyproject.toml \
Expand Down Expand Up @@ -410,7 +418,7 @@ mistral_common[audio]
petit-kernel

# tokenizer extras
fastokens==0.2.0
fastokens==0.3.2
EOT
uv pip install \
-r /tmp/requirements.txt
Expand Down
Loading