Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,16 @@ defaults:

on:
workflow_dispatch:
# `pack/**` is deliberately NOT ignored here or below: it is the only path a
# broken dependency-probe wiring can arrive through, and the guard that catches
# it is tests/pack/test_probe_wiring.py. Ignoring it buys back a few runner
# minutes and gives up the guard entirely.
push:
branches:
- "main"
- "v*-dev"
paths-ignore:
- "!.github/workflows/ci.yml"
- "pack/**"
- "docs/**"
- "tools/**"
- "**.md"
Expand All @@ -36,7 +39,6 @@ on:
- "v*-dev"
paths-ignore:
- "!.github/workflows/ci.yml"
- "pack/**"
- "docs/**"
- "tools/**"
- "**.md"
Expand Down
110 changes: 105 additions & 5 deletions .github/workflows/pack.yml
Original file line number Diff line number Diff line change
Expand Up @@ -221,11 +221,40 @@ jobs:
DOCKER_CONTEXT="${{ github.workspace }}/pack${{ github.event.inputs.post_operation && format('/.post_operation/{0}', github.event.inputs.post_operation) || '' }}/${{ matrix.backend }}"
echo "docker_context=${DOCKER_CONTEXT}" >> $GITHUB_OUTPUT

if [[ -f ${DOCKER_CONTEXT}/Dockerfile.${{ matrix.service }} ]]; then
echo "docker_file=${DOCKER_CONTEXT}/Dockerfile.${{ matrix.service }}" >> $GITHUB_OUTPUT
else
echo "docker_file=${DOCKER_CONTEXT}/Dockerfile" >> $GITHUB_OUTPUT
DOCKER_FILE="${DOCKER_CONTEXT}/Dockerfile"
if [[ -f ${DOCKER_FILE}.${{ matrix.service }} ]]; then
DOCKER_FILE="${DOCKER_FILE}.${{ matrix.service }}"
fi
echo "docker_file=${DOCKER_FILE}" >> $GITHUB_OUTPUT

# Report whether this Dockerfile exposes the dependency export stage.
#
# Only post operations are gated on this. The ones predating dependency
# probing have no `<service>-deps` stage, and exporting it would fail the
# run -- after the image was already pushed. Skipping leaves the recorded
# dependencies stale but never wrong, since `pack/merge_runner.sh` only
# updates the entries it actually probed. The normal path is never gated,
# so a backend Dockerfile that drops the stage still fails loudly.
HAS_DEPENDENCIES_STAGE=false
if grep -qiE "^[[:space:]]*FROM[[:space:]]+.*[[:space:]]+AS[[:space:]]+${{ matrix.service }}-deps([[:space:]]|#|$)" "${DOCKER_FILE}"; then
HAS_DEPENDENCIES_STAGE=true
fi
echo "[INFO]: dependency export stage '${{ matrix.service }}-deps' present in ${DOCKER_FILE}: ${HAS_DEPENDENCIES_STAGE}"
echo "has_dependencies_stage=${HAS_DEPENDENCIES_STAGE}" >> $GITHUB_OUTPUT

# Flatten the whitelist into the build argument consumed by
# pack/shared/probe_dependencies.sh. An empty one must fail the job here:
# the script treats empty as "skip probing", the escape hatch for manual
# local builds, which in CI would silently ship an unprobed image.
DEPENDENCY_PACKAGES="$(jq -er '[.[][]] | join(" ")' "${{ github.workspace }}/pack/dependencies.json")"
# Match on "has a non-whitespace character" rather than stripping spaces:
# `${VAR// /}` strips spaces only, so a tab would pass as non-empty.
if [[ ! "${DEPENDENCY_PACKAGES}" =~ [^[:space:]] ]]; then
echo "[ERROR]: empty dependency whitelist from 'pack/dependencies.json'" >&2
exit 1
fi
echo "[INFO]: dependency packages ($(echo "${DEPENDENCY_PACKAGES}" | wc -w | tr -d ' ')): ${DEPENDENCY_PACKAGES}"
echo "dependency_packages=${DEPENDENCY_PACKAGES}" >> $GITHUB_OUTPUT
- name: Package
timeout-minutes: 360
uses: docker/build-push-action@v6
Expand All @@ -242,16 +271,74 @@ jobs:
sbom: false
context: ${{ steps.metadata.outputs.docker_context }}
file: ${{ steps.metadata.outputs.docker_file }}
# Carries pack/shared/, which every backend Dockerfile mounts to reach
# the probe script. The build context itself stays pack/<backend>/.
build-contexts: |
shared=${{ github.workspace }}/pack/shared
platforms: ${{ matrix.platform }}
target: ${{ matrix.service }}
tags: |
${{ steps.metadata.outputs.tags }}
build-args: |
${{ steps.metadata.outputs.build_args }}
DEPENDENCY_PACKAGES=${{ steps.metadata.outputs.dependency_packages }}
cache-from: |
${{ steps.metadata.outputs.cache_from }}
cache-to: |
${{ steps.metadata.outputs.cache_to }}
# Pull the probed versions out of the build that just ran: a second build of
# the same Dockerfile targeting the `FROM scratch` export stage, because a
# build can only have one target. It runs on the same builder, so it hits the
# Package step's cache instead of re-running the business layers, and needs no
# `docker pull`. It exports no cache -- Package already wrote it.
#
# A dry run runs it as a `check` call, which is the only thing that verifies
# the `<service>-deps` stage exists and is well formed. That produces no file
# and uploads nothing.
- name: Export Dependencies
if: ${{ github.event.inputs.post_operation == '' || steps.metadata.outputs.has_dependencies_stage == 'true' }}
uses: docker/build-push-action@v6
with:
allow: |
network.host
security.insecure
call: ${{ github.event.inputs.dry_run == 'true' && 'check' || 'build' }}
# Must stay byte-identical to the Package step: these land in the LLB
# definition, so a mismatch changes every following RUN's digest, misses
# the cache Package just populated, and exports a dependencies.json
# describing an image that was never pushed. Change one, change both.
ulimit: |
nofile=65536:65536
shm-size: '16G'
push: false
provenance: false
sbom: false
context: ${{ steps.metadata.outputs.docker_context }}
file: ${{ steps.metadata.outputs.docker_file }}
build-contexts: |
shared=${{ github.workspace }}/pack/shared
platforms: ${{ matrix.platform }}
target: ${{ matrix.service }}-deps
build-args: |
${{ steps.metadata.outputs.build_args }}
DEPENDENCY_PACKAGES=${{ steps.metadata.outputs.dependency_packages }}
cache-from: |
${{ steps.metadata.outputs.cache_from }}
outputs: ${{ github.event.inputs.dry_run == 'false' && format('type=local,dest={0}/dependencies', runner.temp) || '' }}
# The artifact name carries the platform tag, unique per (platform, tag)
# within a run: that is the key `merge_runner.sh` relates probes back with.
- name: Upload Dependencies
if: ${{ github.event.inputs.dry_run == 'false' && (github.event.inputs.post_operation == '' || steps.metadata.outputs.has_dependencies_stage == 'true') }}
uses: actions/upload-artifact@v4
with:
name: dependencies-${{ matrix.platform_tag }}
path: ${{ runner.temp }}/dependencies/dependencies.json
# `upload-artifact@v4` answers a duplicate name with a non-retryable 409.
# `build` carries `timeout-minutes: 360`, so "Re-run all jobs" is routine
# here, and without this every re-run leg would fail after pushing.
overwrite: true
if-no-files-found: error
retention-days: 1

# Merge all architecture images into manifest list.
manifest:
Expand Down Expand Up @@ -288,8 +375,14 @@ jobs:
done

# Submit a PR to merge the runner.
#
# A post operation runs through here as well: it mutates a released image in
# place, so the dependency versions recorded for that tag have to follow.
# `for_release == 'true'` stays required, and doubles as the guard for it: a
# post operation is by definition applied to released tags, and without the
# flag every tag would carry the `-dev` suffix and address no released entry.
merge-runner:
if: ${{ github.event.inputs.dry_run == 'false' && github.event.inputs.for_release == 'true' && github.event.inputs.post_operation == '' }}
if: ${{ github.event.inputs.dry_run == 'false' && github.event.inputs.for_release == 'true' }}
needs:
- expand-matrix
- manifest
Expand All @@ -300,9 +393,16 @@ jobs:
with:
fetch-depth: 1
persist-credentials: false
- name: Download Dependencies
uses: actions/download-artifact@v4
with:
pattern: dependencies-*
path: ${{ runner.temp }}/dependencies
- name: Merge Runner
env:
INPUT_BUILD_JOBS: ${{ needs.expand-matrix.outputs.build_jobs }}
INPUT_DEPENDENCIES_DIR: ${{ runner.temp }}/dependencies
INPUT_POST_OPERATION: ${{ github.event.inputs.post_operation }}
run: ${{ github.workspace }}/pack/merge_runner.sh
- name: Generate Pull Request Token
id: generate-prt
Expand Down
149 changes: 148 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ backends.
- [Directory Structure](#directory-structure)
- [Dockerfile Convention](#dockerfile-convention)
- [Docker Image Naming Convention](#docker-image-naming-convention)
- [Dependency Versions](#dependency-versions)
- [Integration Process](#integration-process)

## Onboard Services
Expand Down Expand Up @@ -196,6 +197,36 @@ ENTRYPOINT [ "tini", "--" ]

```

### Example Build Command

Each Dockerfile is built with the backend directory as its build context:

```bash
cd pack/cuda

docker buildx build \
--file Dockerfile.vllm \
--target vllm \
--build-context shared=../shared \
--build-arg DEPENDENCY_PACKAGES="$(jq -er '[.[][]] | join(" ")' ../dependencies.json)" \
--tag gpustack/runner:cuda13.0-vllm0.29.0 \
.
```

Two of these flags are mandatory and one is optional:

- `--target {SERVICE}` is **mandatory**. Every Dockerfile ends with an export-only
`FROM scratch AS {SERVICE}-deps` stage that carries nothing but the probed dependency manifest. Without
`--target`, Docker builds the *last* stage in the file and hands back an empty image.
- `--build-context shared=../shared` is **mandatory** as well. The dependency probe script lives in
[pack/shared](pack/shared) so that all backends share one copy, and the build context of `pack/{BACKEND}/`
cannot reach it with a plain `COPY`; a named build context is the only way in. Omitting the flag makes the
`RUN --mount=type=bind,from=shared` step resolve `shared` as an image name and fail the build.
- `--build-arg DEPENDENCY_PACKAGES=...`, in contrast, is optional and is the escape hatch for manual local
builds. Leaving it empty skips probing, and the resulting image then has **no**
`/etc/gpustack-runner/dependencies.json` — which is also how the data pipeline tells "never probed" apart
from "probed, nothing installed".

## Docker Image Naming Convention

The Docker image naming convention is as follows:
Expand Down Expand Up @@ -225,6 +256,115 @@ The Docker image naming convention is as follows:
`gpustack/runner:cann8.1-910b-vllm0.9.2-dev`.
3. After testing, rename the multi-architecture image to the final tag, e.g. `gpustack/runner:cann8.1-910b-vllm0.9.2`.

## Dependency Versions

Besides the image tag, each entry of [runner.py.json](gpustack_runner/runner.py.json) may carry a
`dependencies` map, which records the versions of a whitelisted set of Python packages **as actually
installed in the built image**, not as declared by the Dockerfile `ARG`s. The whitelist lives in
[pack/dependencies.json](pack/dependencies.json) and the probe runs at build time.

```json
{
"docker_image": "gpustack/runner:cann9.1-a3-vllm0.23.0",
"dependencies": {
"lmcache": "0.4.3",
"lmcache-ascend": "0.4.3",
"ray": "2.54.0",
"torch": "2.10.0",
"torch-npu": "2.10.0rc1",
"vllm-ascend": "0.23.0"
}
}
```

### One Name, Several Distributions

Keys are the dependency names of the whitelist. Most map one to one onto a distribution, but a name may
cover several, highest priority first:

```json
{ "mooncake-transfer-engine": ["mooncake-transfer-engine-npu", "mooncake-transfer-engine-rocm", "mooncake-transfer-engine"] }
```

The probe reports raw distribution names, and `pack/merge_runner.sh` folds them onto the dependency name —
the first distribution of the list that is installed wins — so a consumer asking about
`mooncake-transfer-engine` never has to know the accelerator naming conventions.

Two conditions must both hold before grouping distributions under one name:

1. **They must be mutually exclusive** — at most one of them can be installed in any given image. Folding
keeps a single winner, so grouping distributions that *coexist* silently discards one of them. `torch`
and `torch-npu` look like such a pair by their names, but `torch-npu` pins `torch==<same version>` and is
the NPU backend *on top of* it: both are installed, with different versions that mean different things.
The same holds for `lmcache` and `lmcache-ascend` — a CANN image carries both, and they do not even track
the same version.
2. **Their versions must be comparable** — the same versioning scheme, ideally the same release line. A
grouped name yields one specifier for all of them, so a specifier that is meaningful for one and
meaningless for another gives a confidently wrong answer. `sglang-kernel` and `sgl-kernel-npu` *are*
mutually exclusive, yet they stay separate: the former is `0.4.6.post1` and the latter is a date version
`2026.6.1`, so `>=0.4.5` would match the NPU build for no reason at all.

`mooncake-transfer-engine` satisfies both, which is why it is the one grouped entry: its `-npu`, `-rocm` and
generic builds are one per platform and share a release line (`0.3.11.post1` / `0.3.10.post2`).

When in doubt, give each distribution its own name. That records both facts and asserts nothing.

The raw, unfolded probe result stays inside the image at `/etc/gpustack-runner/dependencies.json`, so
`docker run --rm <image> cat /etc/gpustack-runner/dependencies.json` still shows every distribution and
version for troubleshooting.

### Absent Field vs. Absent Key

The field has two levels of meaning, and conflating them leads to wrong conclusions:

| State | Meaning |
|---------------------------------------|-------------------------------------------------------------------------------------|
| `dependencies` is absent | The image was **never probed** — built before probing existed, or built without the whitelist |
| `dependencies` is a map missing a key | The image **was** probed and the package is **not installed** |

### Querying

All three query entries — `list_runners`, `list_backend_runners` and `list_service_runners` — accept a
`dependencies` argument: a tuple of `(dependency name, PEP 440 specifier)` pairs, matched directly against
the `dependencies` map of each entry.

```python
# Images whose lmcache is new enough.
list_runners(service="vllm", dependencies=(("lmcache", ">=0.4.6"),))

# Multiple conditions.
list_runners(backend="cuda", dependencies=(("lmcache", ">=0.4.6"), ("torch", ">=2.9"),))

# An empty specifier asks only whether the package is installed.
list_runners(backend="cuda", dependencies=(("vllm-omni", ""),))

# Strict mode: drop images that were never probed.
list_runners(
backend="cuda",
dependencies=(("lmcache", ">=0.4.6"),),
with_unknown_dependencies=False,
)
```

Five behaviors to keep in mind:

1. **Conditions are ANDed.** A runner must satisfy every pair to be returned; there is no "any of" form.
2. **Unprobed runners are kept by default.** A runner without a `dependencies` field is never filtered out by
a dependency condition, so that adding a condition does not make every pre-existing image disappear at
once. Pass `with_unknown_dependencies=False` to tighten this to "only runners known to satisfy the
condition" — expect a much shorter list until the fleet has been rebuilt.
3. **Pre-releases match.** Matching is done with `prereleases=True`, because `rc` versions are routine here
(`vllm-ascend 0.20.2rc1`, `sglang 0.5.10rc0`); without it `>=0.20.0` would silently skip `0.20.2rc1`. Note
the converse, which is **not** a bug: `>=0.4.6` does not match `0.4.6rc1`, because PEP 440 orders
`0.4.6rc1 < 0.4.6`. The same rule applies to `dev` versions — `0.27.0rc2.dev25+g<sha>` does not satisfy
`>=0.27.0`. Write the bound you actually mean (`>=0.4.6rc1`) instead of "fixing" the comparison.
4. **An empty specifier is a pure existence check.** `("vllm-omni", "")` matches any image that has the
package, whatever its version. This is the right form for a package installed from a commit, whose
version string carries no information — comparing it would give a confidently wrong answer.
5. **An unknown name matches nothing; it does not raise.** The whitelist is a build-side file and is not
shipped with the library, so there is nothing to check a name against. A misspelled name simply yields an
empty result, which is indistinguishable from "no image qualifies" — callers own their spelling.

## Integration Process

### Ingesting a New Accelerated Backend
Expand All @@ -246,7 +386,14 @@ To add support for a new inference service:
2. Update [pack.yml](.github/workflows/pack.yml) to include the new service in the build matrix.
3. Update [matrix.yml](pack/matrix.yaml) to include the new service.
4. Update `_RE_DOCKER_IMAGE` in [runner.py](gpustack_runner/runner.py) to recognize the new service.
5. [Optional] Update [tests](tests/gpustack_runner) if necessary.
5. Review [pack/dependencies.json](pack/dependencies.json) for the key packages the new service brings
in. A package that is not listed there is never probed for any image, and consumers have no way to ask
about it. The bar for listing one is **it directly decides whether a model or the inference backend
starts** *and* **it is updated often or breaks compatibility** — every entry is recorded for every image,
so the list is meant to stay short. Before grouping accelerator-specific variants under one name, confirm
they are mutually exclusive — see [One Name, Several Distributions](#one-name-several-distributions);
when in doubt, give each its own name, which records both facts and asserts nothing.
6. [Optional] Update [tests](tests/gpustack_runner) if necessary.

## License

Expand Down
Loading
Loading