Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 15 additions & 10 deletions api/worker/v1alpha1/generated.pb.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

13 changes: 12 additions & 1 deletion api/worker/v1alpha1/generated.proto

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

22 changes: 20 additions & 2 deletions api/worker/v1alpha1/kv_cache_backend.go
Original file line number Diff line number Diff line change
Expand Up @@ -339,8 +339,19 @@ type KVCacheBackendLeader struct {
// domain that belongs to whoever typed it. The store's global -quota_bytes flag stays in
// extraArgs for the converse reason: no other API needs to interpret it.
//
// Unset and false both mean no ledger, and unset renders NO flag rather than an explicit false.
MultiTenancy bool `json:"multiTenancy,omitempty" protobuf:"varint,4,opt,name=multiTenancy"`
// IT DEFAULTS TO TRUE, because the ledger is what makes the rest of this API mean what it says:
// without it a KVCachePoolBinding's ceiling is recorded but not enforced, and a master serves one
// reuse domain only, so a second Binding on it is refused. Mooncake has taken the switch since
// 0.3.12, and the default store image is on 0.3.13.post1.
//
// OMITTING THIS KEY AND WRITING `multiTenancy: false` ARE DIFFERENT — the first takes the
// default, the second declines the ledger: only the explicit false renders no switch. A store
// image older than Mooncake 0.3.12 does not recognize the switch and its master exits at
// startup, so a backend on such an image, including the 0.3.10.post2 variants this project
// also publishes, sets false here. Read it through KVCacheBackendLeader.MultiTenancyEnabled.
//
// +k8s:validation:default=true
MultiTenancy *bool `json:"multiTenancy,omitempty" protobuf:"varint,4,opt,name=multiTenancy"`

// ExtraArgs passes flags this API does not enumerate straight through to the leader, after
// the derived ones. Each entry is one flag token of its own, "-flag" or "-flag=value", and the
Expand Down Expand Up @@ -375,6 +386,13 @@ type KVCacheBackendLeader struct {
ExtraEnv []InstanceEnvVar `json:"extraEnv,omitempty" protobuf:"bytes,6,rep,name=extraEnv"`
}

// MultiTenancyEnabled reports whether the leader runs with its tenant ledger. An unset field reads
// as the schema default, true, so an object that never passed the API server, such as one built
// in a test or by a fake client, answers the same as one that did.
func (in KVCacheBackendLeader) MultiTenancyEnabled() bool {
return in.MultiTenancy == nil || *in.MultiTenancy
}

// KVCacheBackendLeaderHighAvailability turns leader election on, and carries how members find the
// leader it elects.
//
Expand Down
6 changes: 5 additions & 1 deletion api/worker/v1alpha1/zz_generated.crds.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

5 changes: 5 additions & 0 deletions api/worker/v1alpha1/zz_generated.deepcopy.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion api/worker/zz_generated.openapi.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

12 changes: 9 additions & 3 deletions docs/kv-cache/backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ spec:
image: docker.io/kvcacheai/mooncake:0.3.13
connection:
managed: # or external: — exactly one
leader: {} # replicas and allocationStrategy default
leader: {} # replicas, allocationStrategy and multiTenancy default
members:
- nodeSelector: {kubernetes.io/os: linux}
medium: DRAM # what this group's SEGMENT is made of: DRAM or VRAM
Expand Down Expand Up @@ -192,6 +192,12 @@ Each variant is built on `0.3.13.post1`, the line vLLM's supported clients are o
minimum](../reference/engine-versions.md). SGLang's clients are on the 0.3.12 line, which this
project does not build. Which line a backend needs is that table's question.

**A `0.3.10.post2` variant also needs `leader.multiTenancy: false` written out.** The tenant ledger
hangs on the master's `-enable_multi_tenants` switch, which Mooncake took in 0.3.12, and the field
defaults on — so on an older image the default renders a flag the master does not recognize, and it
exits at startup. The explicit false renders no flag, which is the command line such an image has
always run.

**A VRAM group needs a build with VRAM segments compiled in (`USE_VRAM_SEGMENT=ON`), and the stock
`-cpu` default is not one.** VRAM segments exist only on the `0.3.13` line — the `0.3.10.post2`
variants carry the vendor transfer engine without them — so a VRAM group always names a
Expand Down Expand Up @@ -235,8 +241,8 @@ write then fails at transfer time with `RPC_FAIL (-900)`.

Two posts of one minor line share their RPC signatures and interoperate. The 0.3.12 and 0.3.13
lines do not: the method names are unchanged, so the client reaches the handler and mis-decodes the
arguments. Multi-tenancy moves neither: with it off — the master's own default — every request
resolves to the default tenant.
arguments. Multi-tenancy moves neither: with it off — a declared `multiTenancy: false`, the field
defaulting on — every request resolves to the default tenant.

**The client's version is a property of the engine image, not of anything on this CR.** Which
client each supported engine's runner image carries, and so which line its store runs, is in the
Expand Down
24 changes: 15 additions & 9 deletions docs/kv-cache/walkthrough.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,11 @@ later. Naming a published upstream image here works until Step 4 and then does n
`capacityPerMember` is charged to each member Pod's host memory request, so it is a claim on the node
and not a hint. One member Pod runs per node the selector matches.

`leader: {}` takes the field defaults, `multiTenancy` included, so this master keeps a per-tenant
quota ledger and Steps 2 and 3 read against it. A backend pinned to a store image from before
Mooncake 0.3.12 is the one exception and declares `multiTenancy: false` out loud — see
[The project's own build variants](backend.md#the-projects-own-build-variants).

Wait for it, and read what it actually says:

```console
Expand Down Expand Up @@ -111,18 +116,19 @@ workloads sharing a domain share cached blocks, so a domain that could be edited
workload start reading blocks another tokenizer wrote. Pick it to match the model and engine
settings the deployments in this namespace will run; a second, different model gets a second Binding.

**`domain.name` is left out, so it is `default`.** The backend from Step 1 runs without
multi-tenancy, so no tenant is forwarded to the engines and the name only records the registration.
On a multi-tenant backend, name each domain; the name is the tenant id the engines are handed.
**`domain.name` is left out, so it is `default`.** The backend from Step 1 runs with multi-tenancy
— `leader.multiTenancy` defaults on — so `default` is the tenant id the engines are handed, the
store's own tenant for a writer that names none. Name each domain once a second Binding shares the
master; the name is what keeps the two apart.

**A quota ceiling is not a reservation.** On a backend with `leader.multiTenancy: true` it is the
most this namespace may hold at once, and going over it does not fail a write — see
**A quota ceiling is not a reservation.** It is the most this namespace may hold at once, and going
over it does not fail a write — see
[What a full quota actually does](pool.md#what-a-full-quota-actually-does).

**On the backend from Step 1 the ceiling is recorded but not enforced.** That leader runs without
multi-tenancy, so the master holds no tenant ledger: the pool is admitted with a warning, the Binding
reports `QuotaGranted=True` with reason `Unenforced` and no `EFFECTIVE` figure, and every write lands
in the store's default tenant. Set `leader.multiTenancy: true` in Step 1 to have ceilings enforced.
**The ceiling is enforced because the Step-1 leader carries its tenant ledger, which the default
gave it.** A backend declared `leader.multiTenancy: false` holds no ledger instead: its pool is
admitted with a warning, the Binding reports `QuotaGranted=True` with reason `Unenforced` and no
`EFFECTIVE` figure, and every write lands in the store's default tenant.

**A multi-tenant master refuses a tenant name absent from its ledger.** An engine that ignores the
injected tenant then needs a second Binding whose domain is `default`, or that leaves `name` out — see
Expand Down

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading