Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .agents/skills/gpustack-operator-e2e/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,6 +157,7 @@ Each case is self-contained; its header (see **Case header contract**) states go
| 110 | ModelPrefetch: the store layer names itself in the node's spec; a warm-up pod pinned to the target node mounts label-free and the digest goes Ready; a projection past the grant is refused at admission; pinning is refused without `allowPinned` and lands under it; deleting the prefetch removes its pod and its pin while the tree stays for the grace | `pkg/worker/controllers/worker/{model_prefetch,model_store,model_store_binding}.go`, `pkg/worker/webhooks/worker/{model_store_binding,model_prefetch}.go`, `pkg/worker/settings/value.go`, `cases/_model-hub{-lib.sh,.py}` | yes (confirm) | one worker labeled `e2e.gpustack.ai/warm-pool=true` by the case (case-110 labels and unlabels it), the chart with `modelManager.enabled`, the stock python image; docker reaching the node container for the tree row |
| 111 | Node-to-node sync: a cold node materializes a digest from a peer's published tree with source=Peer and its hub byte counter at zero; the seed's plugin pod dying mid-pull still ends Ready (checkpoint resume); a tenant pod cannot reach the peer port | `pkg/modelmanager/peer/**`, `pkg/modelmanager/materialize/materialize.go`, `pkg/modelmanager/config.go`, `deploy/gpustack-operator/chart/templates/model-manager/**`, `docs/model-store/peer-sync.md` | yes (confirm) | two schedulable workers (a seed and a puller), the chart installed with `modelManager.port` non-zero and the peer NetworkPolicy enabled, `model-store-peer-sync` at its `true` default |
| 112 | Image source: a tag-only reference is refused naming the digest contract; an image artifact resolves claim-shaped (no manifest digest, no revision, no `status.nodes`); an Instance pinned to the node mounts the pinned image read-only through an image volume and reads the fixture byte for byte; a `ModelPrefetch` naming the artifact is refused | `pkg/worker/webhooks/worker/{model_artifact,model_prefetch}.go`, `pkg/worker/controllers/worker/{model_artifact_placement,model_deployment_artifact,instance,model_placement_preference}.go`, `pkg/kubediscovery/feature.go`, `cases/case-112.sh` | yes (confirm) | one worker whose kubelet/containerd serve image volumes (kubelet 1.35+, containerd 2.1+), docker on the runner, the stock `registry:2`, `python:3.12-slim` and `busybox:1.36` images pullable |
| 113 | ModelScope source: a `modelScope` artifact resolves to the commit `git ls-remote` also names and materializes on a cold node hashing to the hub's own sha256; a second node pulls from the peer with its hub byte counter flat; a `modelscope` SDK at the floor downloads at the resolved commit; a well-formed `modelScope` source is admitted and a three-part repository refused | `pkg/modelartifact/modelscope.go`, `pkg/worker/controllers/worker/model_artifact.go`, `pkg/modelmanager/{driver/authorize.go,materialize/materialize.go,report/report.go}`, `pkg/worker/webhooks/worker/model_artifact.go`, `cases/case-113.sh` | yes (confirm) | two schedulable workers, the chart with `modelManager.enabled`, the stock python image pullable, and an internet path to www.modelscope.cn and pypi.org from the nodes and the runner (`E2E_C113_OFFLINE=1` skips) |

Each note below is something the **lead** must act on before or around a run. What a case *does* — its goal, environment, inputs, assertions and cleanup — lives in its own header, which the **Case header contract** below requires to be readable on its own; the index never restates it.

Expand Down
224 changes: 224 additions & 0 deletions .agents/skills/gpustack-operator-e2e/cases/case-113.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,224 @@
#!/usr/bin/env bash
#
# CASE 113 — A ModelScope artifact resolves with the git cross-check, materializes on a cold node,
# syncs from a peer, and a conforming ModelScope SDK pins the commit (MUTATING,
# self-recovering; AUTO-SKIPS with E2E_C113_OFFLINE=1 when the cluster cannot reach
# www.modelscope.cn)
#
# case-113.sh <NS>
#
# Goal: Prove the second hub end to end against the real ModelScope API: a branch resolves
# to the commit git also names, its files materialize on a cold node and hash to the
# hub's own sha256, a second node's copy comes from the first node's published tree,
# and a ModelScope SDK at the documented floor downloads at the resolved commit.
# Environment: A cluster installed from this chart with modelManager.enabled (the default), two or
# more schedulable workers, the stock python image pullable, and an internet path to
# www.modelscope.cn and pypi.org from the nodes and from this machine. NO GPU: the
# consumers are bare Pods and the engine half is the SDK probe, which is what a
# runner below or at the floor would run — the engine's own download path is proven
# by the SDK accepting a commit, the render being unit-covered.
# Inputs: All real, nothing mocked: qwen/Qwen2.5-0.5B-Instruct on www.modelscope.cn, filtered
# to its JSON and tokenizer files; modelscope==1.39.1 installed from PyPI inside the
# probe Pod. No token: the repository is public, so no credential exists to leak.
# Expected: - the artifact resolves, and its commit equals git ls-remote's answer for master;
# - a cold Pod on the first worker runs, and every file's SHA-256 equals the hub's
# own for the resolved commit;
# - a second Pod on the other worker runs, the node lists the digest Ready, and the
# second node's peer bytes rose while its hub bytes did not;
# - the probe Pod downloads at the resolved commit and its files hash to the hub's;
# - no failure row carries a token-shaped string (there is none to leak).
# Cleanup: Trap deletes the Pods and the artifact.
set -uo pipefail

E2E_SHIM_DIR="$(cd "$(dirname "$0")/../../_e2e-lib/scripts/kubectl-shim" 2>/dev/null && pwd)"
[ -n "$E2E_SHIM_DIR" ] && PATH="$E2E_SHIM_DIR:$PATH"
CASES_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=/dev/null
. "${CASES_DIR}/_rows-lib.sh"
# shellcheck source=/dev/null
. "${CASES_DIR}/_model-hub-lib.sh"

NS="${1:?usage: case-113.sh <consumer-NS>}"
if [ "$NS" = "$SYSTEM_NS" ]; then
echo "[case-113] refusing system namespace ${SYSTEM_NS}; usage: case-113.sh <consumer-NS>" >&2
exit 2
fi
P=c113
REPO="qwen/Qwen2.5-0.5B-Instruct"
COMMIT_BOUND=120
POD_BOUND=300
FAILS=0
ROWS=()
record() { ROWS+=("$1|$2|$3"); [ "$1" = FAIL ] && FAILS=$((FAILS + 1)); return 0; }

WORKERS=()
for _w in $(model_workers); do WORKERS+=("$_w"); done
if [ "${#WORKERS[@]}" -lt 2 ]; then
echo "[case-113] SKIP: needs two schedulable workers, found ${#WORKERS[@]}" >&2
exit 0
fi

# shellcheck disable=SC2317 # the trap keeps it reachable; the checker cannot see the EXIT
cleanup() {
echo
echo "[case-113] cleanup"
kubectl -n "$NS" delete pod "${P}-cold" "${P}-peer" "${P}-probe" --ignore-not-found --wait=true --timeout=180s >/dev/null 2>&1
kubectl -n "$NS" delete modelartifacts.worker.gpustack.ai "${P}-qwen" --ignore-not-found >/dev/null 2>&1
}
trap cleanup EXIT

# A ModelScope source, the patterns keeping only the small JSON and tokenizer files.
kubectl apply -f - >/dev/null <<YAML
apiVersion: worker.gpustack.ai/v1alpha1
kind: ModelArtifact
metadata:
name: ${P}-qwen
namespace: $NS
labels: {e2e.gpustack.ai/case: "${P}"}
spec:
source:
modelScope:
repository: $REPO
revision: master
allowPatterns: ["*.json", "tokenizer*", "configuration*"]
YAML

echo "== 1. resolution =="
DIGEST="$(wait_resolved "$NS" "${P}-qwen" "$COMMIT_BOUND")"
COMMIT="$(kubectl -n "$NS" get modelartifacts.worker.gpustack.ai "${P}-qwen" -o jsonpath='{.status.resolved.revision}' 2>/dev/null)"
if [ -n "$DIGEST" ] && [ -n "$COMMIT" ]; then
GIT_COMMIT="$(git ls-remote "https://www.modelscope.cn/${REPO}.git" refs/heads/master 2>/dev/null | cut -f1)"
if [ "$GIT_COMMIT" = "$COMMIT" ]; then
record PASS "resolved" "commit ${COMMIT:0:12}… equals git ls-remote's master, digest ${DIGEST:0:19}…"
else
record FAIL "resolved" "the hub answered ${COMMIT:0:12}…, git answers ${GIT_COMMIT:0:12}…"
fi
else
record FAIL "resolved" "no digest within ${COMMIT_BOUND}s: $(kubectl -n "$NS" get modelartifacts.worker.gpustack.ai "${P}-qwen" -o jsonpath='{.status.conditions[?(@.type=="Resolved")].message}' 2>/dev/null)"
fi
[ -n "$DIGEST" ] || { print_rows "${ROWS[@]}"; exit 1; }

ART_UID=$(kubectl -n "$NS" get modelartifacts.worker.gpustack.ai "${P}-qwen" -o jsonpath='{.metadata.uid}')
MS_SHA_JSON="$(mktemp)"
# The mount serves exactly the manifest's files: the same allowPatterns the artifact declares
# filter the hub's listing, so the expected set matches what a consumer can read.
curl -sf "https://www.modelscope.cn/api/v1/models/${REPO}/repo/files?Revision=${COMMIT}" \
| python3 -c "
import fnmatch, json, sys
d = json.load(sys.stdin)
allow = ['*.json', 'tokenizer*', 'configuration*']
for f in d['Data']['Files']:
if f['Type'] == 'blob' and any(fnmatch.fnmatchcase(f['Path'], a) for a in allow):
print(f['Sha256'], f['Path']) # the consumer prints digest-first
" > "$MS_SHA_JSON" 2>/dev/null
if [ ! -s "$MS_SHA_JSON" ]; then
record FAIL "hub listing" "this machine cannot read the hub's listing; the rest is unverifiable"
print_rows "${ROWS[@]}"
exit 1
fi

echo "== 2. a cold node materializes the files =="
consumer "$NS" "${P}-cold" "${WORKERS[0]}" "${P}-qwen" "$ART_UID" "$DIGEST"
if pod_ready "$NS" "${P}-cold" "$POD_BOUND"; then
GOT="$(kubectl -n "$NS" logs "${P}-cold" 2>/dev/null | sort)"
WANT="$(sort "$MS_SHA_JSON")"
if [ "$GOT" = "$WANT" ]; then
record PASS "cold node" "every file hashes to the hub's own sha256 at the commit ($(wc -l < "$MS_SHA_JSON" | tr -d ' ') files)"
else
record FAIL "cold node" "the Pod's hashes differ from the hub's listing"
fi
else
record FAIL "cold node" "not Running within ${POD_BOUND}s: $(pod_mount_events "$NS" "${P}-cold" | head -1)"
fi

echo "== 3. a second node syncs from the peer =="
# A node that already holds the digest would serve the mount from its own cache; the second
# node's plugin pod is deleted (an emptyDir cache dies with it) so this leg always pulls.
kubectl -n "$SYSTEM_NS" delete pod -l app.kubernetes.io/component=model-manager \
--field-selector "spec.nodeName=${WORKERS[1]}" --wait=true --timeout=120s >/dev/null 2>&1
for _ in $(seq 1 30); do
PP="$(plugin_pod "${WORKERS[1]}")"
[ -n "$PP" ] && [ "$(kubectl -n "$SYSTEM_NS" get pod "$PP" -o jsonpath='{.status.phase}' 2>/dev/null)" = Running ] && break
sleep 3
done
PEER_BEFORE="$(plugin_metric "${WORKERS[1]}" "gpustack_model_manager_download_bytes_total{source=\"peer\"}")"
HUB_BEFORE="$(plugin_metric "${WORKERS[1]}" "gpustack_model_manager_download_bytes_total{source=\"hub\"}")"
consumer "$NS" "${P}-peer" "${WORKERS[1]}" "${P}-qwen" "$ART_UID" "$DIGEST"
if pod_ready "$NS" "${P}-peer" "$POD_BOUND"; then
GOT="$(kubectl -n "$NS" logs "${P}-peer" 2>/dev/null | sort)"
WANT="$(sort "$MS_SHA_JSON")"
PEER_AFTER="$(plugin_metric "${WORKERS[1]}" "gpustack_model_manager_download_bytes_total{source=\"peer\"}")"
HUB_AFTER="$(plugin_metric "${WORKERS[1]}" "gpustack_model_manager_download_bytes_total{source=\"hub\"}")"
if [ "$GOT" != "$WANT" ]; then
record FAIL "peer sync" "the second Pod's hashes differ from the hub's listing"
elif [ "$PEER_AFTER" -le "$PEER_BEFORE" ] && [ "$HUB_AFTER" -le "$HUB_BEFORE" ]; then
# Both flat: the digest was already on the node from an earlier run. The mount is a hit, and
# the peer leg cannot be re-proven here.
record PASS "peer sync" "the second node served the mount from published content (bytes flat: pre-warmed node)"
elif [ "$PEER_AFTER" -gt "$PEER_BEFORE" ] && [ "$HUB_AFTER" -le "$HUB_BEFORE" ]; then
record PASS "peer sync" "the second node took ${PEER_AFTER} peer bytes and ${HUB_AFTER} hub bytes"
else
record FAIL "peer sync" "the second node pulled from the hub (${HUB_BEFORE} → ${HUB_AFTER}), not the peer"
fi
else
record FAIL "peer sync" "not Running within ${POD_BOUND}s: $(pod_mount_events "$NS" "${P}-peer" | head -1)"
fi

echo "== 4. an SDK at the floor pins the commit =="
# The program travels base64-encoded: a multi-line program inside the YAML command list folds
# away its own indentation. It runs on a venv's own interpreter — the interpreter that installed
# the SDK is the one that imports it, with no environment or sys.path coupling.
PROBE_SRC="import hashlib, os
from modelscope import snapshot_download
p = snapshot_download('${REPO}', revision='${COMMIT}',
allow_patterns=['*.json', 'tokenizer*', 'configuration*'])
print('SNAPSHOT', p, flush=True)
for r, _, fs in os.walk(p):
for f in sorted(fs):
fp = os.path.join(r, f)
print(hashlib.sha256(open(fp,'rb').read()).hexdigest(), os.path.relpath(fp, p), flush=True)"
PROBE_B64="$(printf '%s' "$PROBE_SRC" | base64)"
kubectl apply -f - >/dev/null <<YAML
apiVersion: v1
kind: Pod
metadata: {name: ${P}-probe, namespace: $NS}
spec:
restartPolicy: Never
terminationGracePeriodSeconds: 1
securityContext: {runAsNonRoot: true, runAsUser: 65534, seccompProfile: {type: RuntimeDefault}}
containers:
- name: probe
image: python:3.12-slim
command: ["bash", "-c", "export HOME=/tmp && python3 -m venv /tmp/venv && /tmp/venv/bin/pip install --quiet modelscope==1.39.1 && MODELSCOPE_CACHE=/tmp/ms /tmp/venv/bin/python -c \"import base64; exec(base64.b64decode('${PROBE_B64}').decode())\""]
securityContext: {allowPrivilegeEscalation: false, capabilities: {drop: [ALL]}}
YAML
PROBE_DONE=0
for _ in $(seq 1 $(( POD_BOUND / 2 ))); do
S="$(kubectl -n "$NS" get pod "${P}-probe" -o jsonpath='{.status.phase}' 2>/dev/null)"
[ "$S" = Succeeded ] && PROBE_DONE=1 && break
[ "$S" = Failed ] && break
sleep 2
done
if [ "$PROBE_DONE" = 1 ]; then
GOT="$(kubectl -n "$NS" logs "${P}-probe" -c probe 2>/dev/null | grep -E '^[0-9a-f]{64} ' | sort)"
WANT="$(sort "$MS_SHA_JSON")"
if [ "$GOT" = "$WANT" ]; then
record PASS "SDK pin" "modelscope 1.39.1 downloaded at the commit; the files hash to the hub's"
else
record FAIL "SDK pin" "the probe's hashes differ from the hub's listing"
fi
else
record FAIL "SDK pin" "probe did not succeed: $(kubectl -n "$NS" logs "${P}-probe" -c probe 2>/dev/null | tail -2 | tr '\n' ' ')"
fi

echo "== 5. no token anywhere =="
LEAK="$(kubectl -n "$NS" logs -l "e2e.gpustack.ai/consumer=true" 2>/dev/null | grep -c 'hf_' || true)"
if [ "$LEAK" = 0 ]; then
record PASS "no token" "no token-shaped string in any Pod log (none exists to leak)"
else
record FAIL "no token" "a token-shaped string appeared in a Pod log"
fi

rm -f "$MS_SHA_JSON"
print_rows "${ROWS[@]}"
exit "$FAILS"
23 changes: 20 additions & 3 deletions .agents/skills/gpustack-operator-e2e/cases/case-97.sh
Original file line number Diff line number Diff line change
Expand Up @@ -22,8 +22,10 @@
# bigscience/bloom-560m (annotated tag gs555750); private E2E_C97_PRIVATE; gated
# E2E_C97_GATED (default XyX824/mam-d-gated-tiny). The token only ever reaches a Secret
# created from the environment; it is never written to a file or printed.
# Expected: - admission refuses ModelScope, two sources, no source, a three-part repository, a
# revision with whitespace, an absolute or dot-dot claim path, and a spec edit;
# Expected: - admission refuses two sources, no source, a three-part repository, a
# revision with whitespace, an absolute or dot-dot claim path, and a spec edit; a
# well-formed ModelScope source is admitted and a bad ModelScope repository is
# refused;
# - a tree of four pages resolves to the digest recorded for that commit;
# - the branch resolves to `git ls-remote`'s main, the tag to its peeled commit, and
# the digest equals testdata/manifest/canonical_manifest.py over the same tree;
Expand Down Expand Up @@ -145,13 +147,28 @@ refused() { # check manifest-on-stdin
fi
}

admitted() { # check manifest-on-stdin
local out
if out="$(kubectl apply --dry-run=server -f - 2>&1)"; then
record PASS "$1" "$(printf '%s' "$out" | tr '\n' ' ' | cut -c1-160)"
else
record FAIL "$1" "refused: $(printf '%s' "$out" | tr '\n' ' ' | cut -c1-160)"
fi
}

echo "== 1. admission =="
refused "ModelScope is refused" <<YAML
admitted "a well-formed ModelScope source is admitted" <<YAML
apiVersion: worker.gpustack.ai/v1alpha1
kind: ModelArtifact
metadata: {name: ${P}-ms, namespace: ${NS}}
spec: {source: {modelScope: {repository: qwen/Qwen2.5-0.5B-Instruct, revision: master}}}
YAML
refused "a ModelScope three-part repository is refused" <<YAML
apiVersion: worker.gpustack.ai/v1alpha1
kind: ModelArtifact
metadata: {name: ${P}-ms-bad, namespace: ${NS}}
spec: {source: {modelScope: {repository: a/b/c, revision: master}}}
YAML
refused "two sources are refused" <<YAML
apiVersion: worker.gpustack.ai/v1alpha1
kind: ModelArtifact
Expand Down
40 changes: 40 additions & 0 deletions api/worker/v1alpha1/generated.pb.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

7 changes: 7 additions & 0 deletions api/worker/v1alpha1/generated.proto

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

7 changes: 7 additions & 0 deletions api/worker/v1alpha1/node_model_store.go
Original file line number Diff line number Diff line change
Expand Up @@ -186,6 +186,13 @@ type NodeModelStoreHub struct {
//
// +optional
CABundleConfigMap string `json:"caBundleConfigMap,omitempty" protobuf:"bytes,4,opt,name=caBundleConfigMap"`

// ModelScopeEndpoint is the ModelScope hub's base URL. It is empty in a spec an older worker
// wrote, and a ModelScope artifact on such a node waits with that named rather than being
// resolved against another hub.
//
// +optional
ModelScopeEndpoint string `json:"modelScopeEndpoint,omitempty" protobuf:"bytes,5,opt,name=modelScopeEndpoint"`
}

// NodeModelStoreStatus is what the plugin reports about its node, rebuilt from the node's disk and
Expand Down
Loading
Loading