The decision
On a host-fabric transport (RDMA or EFA), should the engine Pods that consume a pool be granted
the fabric access their members already have — and if so, how is it asked for?
This is the product question left over from #286, which is being closed because the measurement it
was waiting on has been taken. That measurement narrowed the question considerably but does not
answer it: what the access costs is now known; whether to grant it is not.
What is already settled, so it is not re-litigated here
The failure mode is a silent fallback, measured. Both engine-side render branches — the store
connector and the point-to-point connector — are told the fabric protocol, auto-discover the
topology of their own container, find zero HCAs, and install the TCP transport with no error, no
warning and no status condition. The deployment reaches Ready and serves. The object says EFA and
the bytes go over TCP. Verbatim readings are in #286.
The access needed for initialization is one extended resource request, not four things. A
five-cell ablation on hand-built Pods (a controlled experiment, not operator-rendered shapes) removed
one element of the member-side base at a time:
| removed |
fabric initializes? |
| nothing (baseline) |
yes |
hostNetwork |
yes |
| the device resource request |
NO — loud, rc=-1 |
| the device-tree hostPath mount |
yes (that device plugin injects the node itself) |
IPC_LOCK + SYS_RESOURCE |
yes |
Three boundaries travel with that table and must not be dropped when it is quoted:
- The mount's redundancy is a property of that vendor's device plugin, not of the protocol. The
cell that removed the request but kept the mount found the device node visible and
world-writable and the open still refused — the mount makes it seeable, the request makes it
openable.
- Capabilities were removed at initialization only. The probe registers no buffers, and locked
memory limits bite at registration time.
- The probe initializes a transport; it moves no bytes. So "hostNetwork is unnecessary" is
established for initialization, not for a data plane — whether fabric traffic egresses a pod
network namespace is untested.
What has to be decided
Whether. Granting this to engine Pods puts a privilege on tenant-adjacent workloads. The blast
radius is wider than the cluster-scoped member DaemonSet even when the grant is one extended
resource, because the workload that carries it is user-authored. The member-side rule — a privilege
is requested, never inferred — has no obvious engine-side spelling: there is no field on the
engine's owning object that asks for it.
If yes, how. Three shapes, ordered by API surface, carried over from #286 with their costs:
- A. A boolean on the ModelDeployment template, beside the existing
privileged. The user names
the privilege explicitly, which matches the member-side rule. Cost: it splits the fabric contract
in two — the backend says EFA, the user must know to flip this on every engine consuming it, and
nothing refuses the mismatch.
- B. Derive it from the referenced backend. The resolve webhook already fetches the
KVCacheBackend at mutate time, so a ModelDeployment binding a fabric backend could have the same
access rendered into its engine container. One declaration, both sides consistent. Cost: the
privilege becomes inferred from a binding rather than requested on the workload — the rule turns
into "binding a fabric backend is the request", which needs an admission-time guard to stay
honest. It also cannot stand alone for an engine consuming a remote store with no local binding.
- C. A pool-level opt-in. The pool that admits fabric members marks itself fabric-capable and
engines admitted against it inherit the access. Cost: the placement contract (per-role
instanceType) becomes a security contract, which it was never meant to be.
A fourth answer is legitimate and is not the absence of a decision: members only, engines stay on
TCP. If that is the answer it has to be written as a rule, with the consequence stated — see the
sibling issue on pool-side reuse, which is what that rule costs.
Note that the field shape #286 assumed for option B no longer exists: the member's fabric device
resource is no longer a declared field but is derived from the transport protocol, so "read
deviceResourceName the same way the member rendering does" now means "derive it the same way".
What does NOT close this
- Rendering the access without deciding
whether. The ablation says what it costs, not that it
should be granted.
- Documenting the gap. It is documented on both pages a reader would look for it. A deployment
that silently serves at TCP speed is not helped by a page the operator's author read.
- A green run on a cluster where the engines do get the access. What has judgement power is what
happens to a deployment whose engines do not — today that is a healthy-looking deployment, which
is the thing to fix or to declare intended.
- Choosing a shape without saying what refuses a mismatch. All three shapes admit a state where
the backend declares a fabric and an engine does not carry the access. Whichever is chosen has to
say whether that state is refused, reported, or allowed.
/kind enhancement
/area kv-cache
The decision
On a host-fabric transport (
RDMAorEFA), should the engine Pods that consume a pool be grantedthe fabric access their members already have — and if so, how is it asked for?
This is the product question left over from #286, which is being closed because the measurement it
was waiting on has been taken. That measurement narrowed the question considerably but does not
answer it: what the access costs is now known; whether to grant it is not.
What is already settled, so it is not re-litigated here
The failure mode is a silent fallback, measured. Both engine-side render branches — the store
connector and the point-to-point connector — are told the fabric protocol, auto-discover the
topology of their own container, find zero HCAs, and install the TCP transport with no error, no
warning and no status condition. The deployment reaches Ready and serves. The object says
EFAandthe bytes go over TCP. Verbatim readings are in #286.
The access needed for initialization is one extended resource request, not four things. A
five-cell ablation on hand-built Pods (a controlled experiment, not operator-rendered shapes) removed
one element of the member-side base at a time:
hostNetworkIPC_LOCK+SYS_RESOURCEThree boundaries travel with that table and must not be dropped when it is quoted:
cell that removed the request but kept the mount found the device node visible and
world-writable and the open still refused — the mount makes it seeable, the request makes it
openable.
memory limits bite at registration time.
established for initialization, not for a data plane — whether fabric traffic egresses a pod
network namespace is untested.
What has to be decided
Whether. Granting this to engine Pods puts a privilege on tenant-adjacent workloads. The blast
radius is wider than the cluster-scoped member DaemonSet even when the grant is one extended
resource, because the workload that carries it is user-authored. The member-side rule — a privilege
is requested, never inferred — has no obvious engine-side spelling: there is no field on the
engine's owning object that asks for it.
If yes, how. Three shapes, ordered by API surface, carried over from #286 with their costs:
privileged. The user namesthe privilege explicitly, which matches the member-side rule. Cost: it splits the fabric contract
in two — the backend says
EFA, the user must know to flip this on every engine consuming it, andnothing refuses the mismatch.
KVCacheBackendat mutate time, so a ModelDeployment binding a fabric backend could have the sameaccess rendered into its engine container. One declaration, both sides consistent. Cost: the
privilege becomes inferred from a binding rather than requested on the workload — the rule turns
into "binding a fabric backend is the request", which needs an admission-time guard to stay
honest. It also cannot stand alone for an engine consuming a remote store with no local binding.
engines admitted against it inherit the access. Cost: the placement contract (per-role
instanceType) becomes a security contract, which it was never meant to be.A fourth answer is legitimate and is not the absence of a decision: members only, engines stay on
TCP. If that is the answer it has to be written as a rule, with the consequence stated — see the
sibling issue on pool-side reuse, which is what that rule costs.
Note that the field shape #286 assumed for option B no longer exists: the member's fabric device
resource is no longer a declared field but is derived from the transport protocol, so "read
deviceResourceNamethe same way the member rendering does" now means "derive it the same way".What does NOT close this
whether. The ablation says what it costs, not that itshould be granted.
that silently serves at TCP speed is not helped by a page the operator's author read.
happens to a deployment whose engines do not — today that is a healthy-looking deployment, which
is the thing to fix or to declare intended.
the backend declares a fabric and an engine does not carry the access. Whichever is chosen has to
say whether that state is refused, reported, or allowed.
/kind enhancement
/area kv-cache