feat: upgrade vllm from 0.29.0 to 0.30.0 for CUDA and ROCm backends - #280
Open
axsapronov wants to merge 1 commit into
Open
axsapronov wants to merge 1 commit into
axsapronov wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Code Review
This pull request updates vLLM to v0.30.0, SGLang's LMCache to v0.5.5, and upgrades associated dependencies like PyTorch, torchvision, and fastokens across CUDA and ROCm Dockerfiles and the build matrix. To support LMCache's build-time gRPC stub generation, the installation of grpcio and grpcio-tools was added. However, the reviewer identified a critical issue where the specified version 1.78.0 for these gRPC packages does not exist on PyPI, which will cause the image builds to fail. It is recommended to correct this version to a valid release, such as 1.68.0.
axsapronov
force-pushed
the
feature/vllm-0.30
branch
2 times, most recently
from
September 23, 2026 03:47
fcdf988 to
9ec718a
Compare
…30.0rc1, lmcache to 0.5.5 Why these move together: - vllm / omni / torch are one resolved dependency set per image. omni v0.30.0rc1 is rebased onto vLLM 0.30.0, so it cannot pair with a 0.29.0 base; and the 0.30.0 base's prebuilt wheels (vllm, flashinfer, fa3-fwd) are compiled against torch 2.13.0. Taking torch 2.14.0 on top of them is deliberate, riding torch's stable ABI within the major -- the lmcache ABI check in the build and the CI smoke tests are what enforce it. - lmcache moves in lockstep across CUDA/ROCm vLLM and both SGLang files: the MP wire protocol has no version handshake, so 0.5.5 on one side and 0.5.4 on the other is an unsupported pairing. 0.5.5 also generates its gRPC stubs at build time, hence the pinned grpcio/grpcio-tools in every lmcache build stage. What moves: - CUDA cu130: vllm 0.30.0, CUDA 13.0.3, torch 2.13.0 -> 2.14.0+cu130 (new self-conditional "Upgrade Torch" step), torchvision 0.29.0, omni v0.30.0rc1, lmcache 0.5.5, fastokens 0.3.2 - CUDA cu130 lmcache build: compile lmcache's C++ sources with C++20 -- torch 2.14's headers require it (#error in torch/all.h) and lmcache 0.5.5 hardcodes -std=c++17 for g++; nvcc already defaults to C++20 under CUDA 13 - CUDA cu129: vllm 0.30.0 (the v0.30.0-cu129 base exists), stays on the base's torch 2.13.0 -- 2.14.0 is published for cu130 only - ROCm: vllm 0.30.0, omni v0.30.0rc1, keeps the base's torch 2.12.x - SGLang (CUDA + ROCm): lmcache 0.5.5 lockstep - CANN and the SGLang engine version are untouched - README: 0.30.0 added to the CUDA 13.0/12.9 and ROCm 7.2 supported version tables
axsapronov
force-pushed
the
feature/vllm-0.30
branch
from
September 23, 2026 12:58
9ec718a to
87bb1ce
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upgrade vLLM from 0.29.0 to 0.30.0 across the CUDA 13.0, CUDA 12.9, and ROCm 7.2 backends, and pull in the matching torch 2.14.0 (CUDA 13.0), vllm-omni v0.30.0rc1, lmcache 0.5.5 (vLLM and SGLang), and fastokens 0.3.2.
Follows the same pattern as #251 and #261.
Why these move together
What moves
CANN and the SGLang engine version are untouched. README: 0.30.0 added to the CUDA 13.0/12.9 and ROCm 7.2 supported version tables.
Build fix: lmcache C++20 on CUDA 13.0
torch 2.14's headers require a C++20 compiler (
#error C++20 or later compatible compiler is required to use PyTorchintorch/all.h), but lmcache 0.5.5 hardcodes-std=c++17for its g++ sources (setup_extensions/common_cpp.py,setup_extensions/build_profiles/cuda.py). Under the torch 2.14.0 base, thepybind.cpptranslation unit fails to compile against the torch headers. The CUDA 13.0 Dockerfile now rewrites that flag to-std=c++20before building lmcache; nvcc already defaults to C++20 under CUDA 13, so only the g++ path needed the bump. The cu129 base (torch 2.13.0) and the SGLang/ROCm bases (torch 2.13.0 / 2.12.x) are unaffected by C++20 and keep building as before.Upstream references
Patch changes
None. The cuda and rocm patch trees are unchanged — v0.30.0 keeps the refactored
get_open_port()and the upstream shm-broadcast port fix, so001_wrong_dp_ray.patchapplies as before.runner.py.jsonand test fixtures are left for the GHA pack workflow's merge-runner job to regenerate.