Skip to content

feat: upgrade vllm from 0.29.0 to 0.30.0 for CUDA and ROCm backends - #280

Open
axsapronov wants to merge 1 commit into
gpustack:mainfrom
axsapronov:feature/vllm-0.30
Open

axsapronov wants to merge 1 commit into
gpustack:mainfrom
axsapronov:feature/vllm-0.30

Conversation

@axsapronov

@axsapronov axsapronov commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Upgrade vLLM from 0.29.0 to 0.30.0 across the CUDA 13.0, CUDA 12.9, and ROCm 7.2 backends, and pull in the matching torch 2.14.0 (CUDA 13.0), vllm-omni v0.30.0rc1, lmcache 0.5.5 (vLLM and SGLang), and fastokens 0.3.2.

Follows the same pattern as #251 and #261.

Why these move together

  • vllm / omni / torch are one resolved dependency set per image. omni v0.30.0rc1 is rebased onto vLLM 0.30.0, so it cannot pair with a 0.29.0 base; the 0.30.0 base's prebuilt wheels (vllm, flashinfer, fa3-fwd) are compiled against torch 2.13.0, and taking torch 2.14.0 on top of them is deliberate — it rides torch's stable ABI within the major version, enforced by the lmcache ABI check in the build and the CI smoke tests.
  • lmcache moves in lockstep across the CUDA/ROCm vLLM and both SGLang images: the MP wire protocol has no version handshake, so 0.5.5 on one side and 0.5.4 on the other is an unsupported pairing. 0.5.5 also generates its gRPC stubs at build time, hence the pinned grpcio/grpcio-tools 1.84.0 in every lmcache build stage.

What moves

Backend Changes
CUDA 13.0 vllm 0.30.0, CUDA 13.0.3, torch 2.13.0 → 2.14.0+cu130 (new self-conditional "Upgrade Torch" step), torchvision 0.29.0, omni v0.30.0rc1, lmcache 0.5.5, fastokens 0.3.2, lmcache C++ built with C++20 (see below)
CUDA 12.9 vllm 0.30.0 (the v0.30.0-cu129 base exists); stays on the base's torch 2.13.0 — 2.14.0 is published for cu130 only
ROCm 7.2 vllm 0.30.0, omni v0.30.0rc1; keeps the base's torch 2.12.x (the 0.30.0 ROCm base is built from the ROCm pytorch release/2.12 branch)
SGLang (CUDA + ROCm) lmcache 0.5.5 lockstep

CANN and the SGLang engine version are untouched. README: 0.30.0 added to the CUDA 13.0/12.9 and ROCm 7.2 supported version tables.

Build fix: lmcache C++20 on CUDA 13.0

torch 2.14's headers require a C++20 compiler (#error C++20 or later compatible compiler is required to use PyTorch in torch/all.h), but lmcache 0.5.5 hardcodes -std=c++17 for its g++ sources (setup_extensions/common_cpp.py, setup_extensions/build_profiles/cuda.py). Under the torch 2.14.0 base, the pybind.cpp translation unit fails to compile against the torch headers. The CUDA 13.0 Dockerfile now rewrites that flag to -std=c++20 before building lmcache; nvcc already defaults to C++20 under CUDA 13, so only the g++ path needed the bump. The cu129 base (torch 2.13.0) and the SGLang/ROCm bases (torch 2.13.0 / 2.12.x) are unaffected by C++20 and keep building as before.

Upstream references

Patch changes

None. The cuda and rocm patch trees are unchanged — v0.30.0 keeps the refactored get_open_port() and the upstream shm-broadcast port fix, so 001_wrong_dp_ray.patch applies as before. runner.py.json and test fixtures are left for the GHA pack workflow's merge-runner job to regenerate.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates vLLM to v0.30.0, SGLang's LMCache to v0.5.5, and upgrades associated dependencies like PyTorch, torchvision, and fastokens across CUDA and ROCm Dockerfiles and the build matrix. To support LMCache's build-time gRPC stub generation, the installation of grpcio and grpcio-tools was added. However, the reviewer identified a critical issue where the specified version 1.78.0 for these gRPC packages does not exist on PyPI, which will cause the image builds to fail. It is recommended to correct this version to a valid release, such as 1.68.0.

Comment thread pack/cuda/Dockerfile.sglang Outdated
Comment thread pack/cuda/Dockerfile.vllm Outdated
Comment thread pack/rocm/Dockerfile.sglang Outdated
Comment thread pack/rocm/Dockerfile.vllm Outdated
@axsapronov
axsapronov force-pushed the feature/vllm-0.30 branch 2 times, most recently from fcdf988 to 9ec718a Compare September 23, 2026 03:47
…30.0rc1, lmcache to 0.5.5

Why these move together:

- vllm / omni / torch are one resolved dependency set per image. omni
  v0.30.0rc1 is rebased onto vLLM 0.30.0, so it cannot pair with a 0.29.0
  base; and the 0.30.0 base's prebuilt wheels (vllm, flashinfer, fa3-fwd)
  are compiled against torch 2.13.0. Taking torch 2.14.0 on top of them is
  deliberate, riding torch's stable ABI within the major -- the lmcache ABI
  check in the build and the CI smoke tests are what enforce it.
- lmcache moves in lockstep across CUDA/ROCm vLLM and both SGLang files:
  the MP wire protocol has no version handshake, so 0.5.5 on one side and
  0.5.4 on the other is an unsupported pairing. 0.5.5 also generates its
  gRPC stubs at build time, hence the pinned grpcio/grpcio-tools in every
  lmcache build stage.

What moves:
- CUDA cu130: vllm 0.30.0, CUDA 13.0.3, torch 2.13.0 -> 2.14.0+cu130
  (new self-conditional "Upgrade Torch" step), torchvision 0.29.0,
  omni v0.30.0rc1, lmcache 0.5.5, fastokens 0.3.2
- CUDA cu130 lmcache build: compile lmcache's C++ sources with C++20 --
  torch 2.14's headers require it (#error in torch/all.h) and lmcache 0.5.5
  hardcodes -std=c++17 for g++; nvcc already defaults to C++20 under CUDA 13
- CUDA cu129: vllm 0.30.0 (the v0.30.0-cu129 base exists), stays on the
  base's torch 2.13.0 -- 2.14.0 is published for cu130 only
- ROCm: vllm 0.30.0, omni v0.30.0rc1, keeps the base's torch 2.12.x
- SGLang (CUDA + ROCm): lmcache 0.5.5 lockstep
- CANN and the SGLang engine version are untouched
- README: 0.30.0 added to the CUDA 13.0/12.9 and ROCm 7.2 supported
  version tables
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant