Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .agents/skills/gpustack-operator-docs/references/page-map.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,8 +124,8 @@ in one clause and does not describe it.
**Owns** — the leader process end to end: the Deployment and ClusterIP Service and their two ports,
the replica ceiling and the clamp that survives a missing webhook, the update strategy per replica
count, the two probes and why they take different paths, the health document's four fields, and all
of `leader.highAvailability` -- the Lease, the image both roles need, the per-role ServiceAccounts,
and how a missing grant fails on each side.
of leader election -- `leader.electionBackend`, `leader.memberAddressing`, the Lease, the image both
roles need, the per-role ServiceAccounts, and how a missing grant fails on each side.

**Never** — the member groups, the transport, the status algebra. Those stay on `backend.md`, which
links here. It was split out when that page hit both the line and the `##` cap.
Expand Down
4 changes: 2 additions & 2 deletions .agents/skills/gpustack-operator-e2e/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,8 +119,8 @@ Each case is self-contained; its header (see **Case header contract**) states go
| 70 | A routed P/D deployment owns and garbage-collects all six router objects, converges both `spec.router` transitions, remains Starting without an accelerator, and reports every cluster-observable `KVEventsPublishing` status/reason pair | `pkg/worker/controllers/worker/model_deployment_router.go`, `model_deployment_kv_events.go`, `model_deployment_status.go` | yes (confirm) | Single Ready node with no accelerator; CASE 1 has materialized one usable general InstanceType, and the cluster can pull the Mooncake fixture image so a real Ready Binding can make `Publishing` reachable. The router image is deliberately unpullable because readiness is isolated in CASE 71 |
| 71 | A Ready router becomes `status.endpoint`; inability to pull the upstream router image is a stated SKIP rather than loss of CASE 70's lifecycle coverage | `pkg/worker/controllers/worker/model_deployment_status.go`, the default llm-d-router image contract | yes (confirm) | As CASE 70 plus pull access to the default llm-d-router and Envoy images; AUTO-SKIP only for an image-pull reason |
| 73 | An engine under KV turnover writes the shared store, and the other replica's replay of the same prefixes is the reuse the chain exists for — the write half is a guard, the read half a KNOWN-FAILURE DETECTOR pair (case-67 polarity) that FAILS the day cross-replica reuse starts working, and must then be inverted into positive guards | `pkg/worker/controllers/worker/model_deployment_connector.go`, `pkg/worker/controllers/worker/model_deployment_binding.go`, `pkg/worker/kvcache/inject/**`, the engine image pin | yes (confirm) | A real accelerator pool with at least TWO free exclusive cards, model weights hostPath-staged on the accelerator nodes (`E2E_VB_WEIGHTS`, default `/mnt/kvcache-weights`), `E2E_VB_INSTANCE_TYPE` naming the accelerated InstanceType (exit 2 without it), and a registry the cluster can pull the Mooncake image from — a CUDA-only tag crashes members on CPU-only nodes, so the image must be CPU-capable. Where no two cards are free, `E2E_VB_EXISTING_BACKEND` + `E2E_VB_EXISTING_DOMAIN` + `E2E_VB_EXISTING_SERVICES` together run every verdict row against an existing engine pair and create nothing. The first case in the suite to run real vLLM engines; every verdict rides on counters (`master_key_count`, `vllm:external_prefix_cache_hits_total`, `mem_cache_hit_nums_`), never on TTFT |
| 74 | `leader.highAvailability.snapshot` is not in the installed schema — a strict create is refused naming it as an unknown field, a lenient one is accepted with the block pruned — and the store's snapshot flags are refused in leader.extraArgs on create and on an update to a running backend, each refusal on the entry with its reason (a restore can serve another key's bytes; the other keys are read only under a refused switch); the manifest without them, and an image-only update, are accepted as the positive baseline | `api/worker/v1alpha1/kv_cache_backend.go`, `pkg/worker/kvcache/mooncake/keys.go`, `pkg/worker/webhooks/worker/kv_cache_backend.go` | yes (confirm) | any (no GPU, no RDMA, no storage class); the creates and the updates are server-side dry runs, and the one backend the case persists selects no node, so it renders a leader and no member Pod |
| 75 | An image bump on a replicated leader emits the `KVCacheLeaderHandover` event — read in namespace `default`, where a cluster-scoped backend's events land — with its cumulative count equal to the lease's `leaseTransitions`, every leader pod a replacement on the new image, exactly one ready, and the backend settling Ready with every health condition True (`PoolWrites` reports write activity, not health, and reads Unknown on this idle backend); the only mutation is the spec patch, no pod is deleted by hand | `pkg/worker/controllers/worker/kv_cache_backend_handover.go` | yes (confirm) | any (no GPU, no RDMA) + a registry the cluster can pull BOTH pinned Mooncake tags from (`E2E_MOONCAKE_IMAGE` start / `E2E_MOONCAKE_ROLLOUT_IMAGE` target; the pair must each parse this operator's argv, carry the lease backend, and run on CPU-only nodes) |
| 74 | The store's snapshot flags are refused in leader.extraArgs on create and on an update to a running backend, each refusal on the entry with its reason (a restore can serve another key's bytes; the other keys are read only under a refused switch); the manifest without them, and an image-only update, are accepted as the positive baseline | `pkg/worker/kvcache/mooncake/keys.go`, `pkg/worker/webhooks/worker/kv_cache_backend.go` | yes (confirm) | any (no GPU, no RDMA, no storage class); the creates and the updates are server-side dry runs, and the one backend the case persists selects no node, so it renders a leader and no member Pod |
| 75 | The default single leader holds a Lease; `electionBackend: None` with three replicas is refused; scaling from one to three preserves the first leader Pod; an image bump then emits `KVCacheLeaderHandover` with its count equal to the Lease transitions, replaces all leader Pods, and settles Ready with one serving replica | `pkg/worker/controllers/worker/kv_cache_backend_handover.go`, `pkg/worker/kvcache/mooncake/leader_workload.go` | yes (confirm) | any (no GPU, no RDMA) + a registry the cluster can pull BOTH pinned Mooncake tags from (`E2E_MOONCAKE_IMAGE` start / `E2E_MOONCAKE_ROLLOUT_IMAGE` target; both must parse this operator's argv, carry the lease backend, and run on CPU-only nodes) |
| 76 | `RolloutComplete` stays truthful through a second mid-update (every sample is `Unknown/UpdateNotObserved`, `False/Progressing` or `True/Complete`, never a deadline-ish stall), the update converges with the election gate intact, and the member-re-registration dip clears within its window; the case never gates on `kubectl rollout status`, which times out on every multi-replica leader rollout by construction | `pkg/worker/controllers/worker/kv_cache_backend_rollout.go` | yes (confirm) | as CASE 75 (the same image pair and clauses) |
| 77 | The multi-tenant ledger gate: an unregistered tenant's put is refused `-1701` while the client itself stays healthy, and a Pool+Binding whose `domain.name` is the tenant id admits the identical put; the tenant rides the keyword `tenant_id=` (the next positional slot is a TransferEngine pointer and raises), and the teardown drains the domain because a held domain blocks pool deletion open-ended | `pkg/worker/controllers/worker/kv_cache_pool.go` (the domain registration pass), `pkg/worker/kvcache/mooncake/**` (the master argv and lease render) | yes (confirm) | any (no GPU, no RDMA) + a registry the cluster can pull the Mooncake image from (`E2E_MOONCAKE_IMAGE`, CPU-capable, carrying the python client); the backend must be a replicated HA leader — the k8s:// master address and the probe's member Role exist only above one replica |
| 78 | Scaling a role moves its queue's admitted quota by exactly one replica in each direction, and touches no replica it did not add or remove | `pkg/worker/controllers/worker/model_deployment{,_pod_group,_rollout}.go`, `api/worker/v1alpha1/model_deployment.go` | yes (confirm) | any (no GPU) + an InstanceType, the pool's entrance LocalQueue in `<NS>`, and room in the pool for three replicas of one role. Optionally `E2E_MD_INSTANCE_TYPE` / `E2E_MD_IMAGE` / `E2E_MD_SETTLE` |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
members:
- nodeSelector: {kubernetes.io/os: linux}
Expand Down
3 changes: 3 additions & 0 deletions .agents/skills/gpustack-operator-e2e/cases/case-43.sh
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
members:
- nodeSelector: {kubernetes.io/os: linux}
Expand Down Expand Up @@ -585,6 +586,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
members:
- nodeSelector: {gpustack.ai/kvc-e2e-absent: "true"}
Expand Down Expand Up @@ -808,6 +810,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
members:
- nodeSelector: {gpustack.ai/kvc-e2e-absent: "true"}
Expand Down
1 change: 1 addition & 0 deletions .agents/skills/gpustack-operator-e2e/cases/case-44.sh
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
# This case times a lease lapse, so it pins the lease rather than inheriting whatever default
# the operator renders. At the operator's own five minutes the lapse below would have to sleep
Expand Down
5 changes: 2 additions & 3 deletions .agents/skills/gpustack-operator-e2e/cases/case-62.sh
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
# cluster-scoped; the Deployment, Service, Lease, ServiceAccounts, Roles and RoleBindings it renders
# all live in <NS>.
#
# Goal: With leader.replicas=3 and leader.highAvailability={}, three leader processes run and
# Goal: With leader.replicas=3 and leader.electionBackend=Kubernetes, three leader processes run and
# exactly one serves, elected through a Kubernetes Lease. This proves on a real API
# server what rendered objects and unit tests cannot:
# (1) THE STEADY STATE IS AN EQUALITY: 3 desired / 1 ready. Three ready would mean
Expand Down Expand Up @@ -40,7 +40,7 @@
# the wrong image, not a flake. Override with E2E_MOONCAKE_IMAGE; the default below
# carries the backend.
#
# Inputs: All real, nothing mocked. One KVCacheBackend (replicas 3, highAvailability, one DRAM
# Inputs: All real, nothing mocked. One KVCacheBackend (replicas 3, Kubernetes election, one DRAM
# member group of 2Gi per node). The failover is induced by deleting the ready leader
# Pod -- the one the Lease names.
#
Expand Down Expand Up @@ -249,7 +249,6 @@ spec:
managed:
leader:
replicas: 3
highAvailability: {}
members:
- nodeSelector: {kubernetes.io/os: linux}
medium: DRAM
Expand Down
5 changes: 2 additions & 3 deletions .agents/skills/gpustack-operator-e2e/cases/case-63.sh
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
# cluster-scoped; its rendered objects live in <NS>, and the two probe Pods live in a namespace this
# case creates and removes.
#
# Goal: Under leader.highAvailability a member can be handed MOONCAKE_MASTER in one of two
# Goal: Under Kubernetes leader election a member can be handed MOONCAKE_MASTER in one of two
# forms. Path A is "k8s://<ns>/<lease>": the client reads the Lease itself and follows
# view changes, which costs the member's image a compiled-in lease backend and so
# excludes every vendor image. Path B, the default, is the plain leader Service
Expand Down Expand Up @@ -44,7 +44,7 @@
# what proves the probe itself can work before any failover is measured, so a
# compiled-out backend fails there, loudly, rather than as a zero measurement.
#
# Inputs: All real, nothing mocked. One KVCacheBackend (leader replicas 3, highAvailability,
# Inputs: All real, nothing mocked. One KVCacheBackend (leader replicas 3, Kubernetes election,
# one DRAM member group of 2Gi). Two probe Pods running the store image's python
# client: probe A set up with "k8s://<ns>/<backend>-leader" (it gets a ServiceAccount
# bound to the member Role -- `get leases` only -- because that read is exactly what
Expand Down Expand Up @@ -276,7 +276,6 @@ spec:
managed:
leader:
replicas: 3
highAvailability: {}
members:
- nodeSelector: {kubernetes.io/os: linux}
medium: DRAM
Expand Down
3 changes: 1 addition & 2 deletions .agents/skills/gpustack-operator-e2e/cases/case-64.sh
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@
# ready" wait times out. That timeout is the signature of the wrong image, not a
# flake. Override with E2E_MOONCAKE_IMAGE; the default below carries the backend.
#
# Inputs: All real, nothing mocked. One KVCacheBackend (replicas 3, highAvailability, one
# Inputs: All real, nothing mocked. One KVCacheBackend (replicas 3, Kubernetes election, one
# DRAM member group of 2Gi per node). The failure is injected with
# `kubectl delete pod --force --grace-period=0` on the Lease holder: the API object
# vanishes immediately and the container runtime SIGKILLs the process, so the
Expand Down Expand Up @@ -212,7 +212,6 @@ spec:
managed:
leader:
replicas: 3
highAvailability: {}
members:
- nodeSelector: {kubernetes.io/os: linux}
medium: DRAM
Expand Down
6 changes: 3 additions & 3 deletions .agents/skills/gpustack-operator-e2e/cases/case-65.sh
Original file line number Diff line number Diff line change
Expand Up @@ -271,11 +271,11 @@ spec:
protocol: TCP
connection:
managed:
# Required by the schema, and empty is the shape this case wants: one leader process, no
# election. Omitting the key is refused at apply, which the host-directory gate above used to
# Required by the schema; this case uses one leader process without election.
# Omitting the key is refused at apply, which the host-directory gate above used to
# hide -- that gate exits 0, so on any cluster without the directory this case reported
# nothing rather than reporting that it could not build its own fixture.
leader: {}
leader: {electionBackend: None}
members:
- nodeSelector: {kubernetes.io/os: linux}
medium: DRAM
Expand Down
1 change: 1 addition & 0 deletions .agents/skills/gpustack-operator-e2e/cases/case-73.sh
Original file line number Diff line number Diff line change
Expand Up @@ -234,6 +234,7 @@ spec:
connection:
managed:
leader:
electionBackend: None
multiTenancy: true
members:
- nodeSelector: {kubernetes.io/os: linux}
Expand Down
Loading
Loading