Skip to content

enhancement: decide whether engine Pods get host-fabric access, and how it is asked for #511

Description

@thxCode

The decision

On a host-fabric transport (RDMA or EFA), should the engine Pods that consume a pool be granted
the fabric access their members already have — and if so, how is it asked for?

This is the product question left over from #286, which is being closed because the measurement it
was waiting on has been taken. That measurement narrowed the question considerably but does not
answer it: what the access costs is now known; whether to grant it is not.

What is already settled, so it is not re-litigated here

The failure mode is a silent fallback, measured. Both engine-side render branches — the store
connector and the point-to-point connector — are told the fabric protocol, auto-discover the
topology of their own container, find zero HCAs, and install the TCP transport with no error, no
warning and no status condition. The deployment reaches Ready and serves. The object says EFA and
the bytes go over TCP. Verbatim readings are in #286.

The access needed for initialization is one extended resource request, not four things. A
five-cell ablation on hand-built Pods (a controlled experiment, not operator-rendered shapes) removed
one element of the member-side base at a time:

removed fabric initializes?
nothing (baseline) yes
hostNetwork yes
the device resource request NO — loud, rc=-1
the device-tree hostPath mount yes (that device plugin injects the node itself)
IPC_LOCK + SYS_RESOURCE yes

Three boundaries travel with that table and must not be dropped when it is quoted:

  1. The mount's redundancy is a property of that vendor's device plugin, not of the protocol. The
    cell that removed the request but kept the mount found the device node visible and
    world-writable and the open still refused — the mount makes it seeable, the request makes it
    openable
    .
  2. Capabilities were removed at initialization only. The probe registers no buffers, and locked
    memory limits bite at registration time.
  3. The probe initializes a transport; it moves no bytes. So "hostNetwork is unnecessary" is
    established for initialization, not for a data plane — whether fabric traffic egresses a pod
    network namespace is untested.

What has to be decided

Whether. Granting this to engine Pods puts a privilege on tenant-adjacent workloads. The blast
radius is wider than the cluster-scoped member DaemonSet even when the grant is one extended
resource, because the workload that carries it is user-authored. The member-side rule — a privilege
is requested, never inferred
— has no obvious engine-side spelling: there is no field on the
engine's owning object that asks for it.

If yes, how. Three shapes, ordered by API surface, carried over from #286 with their costs:

  • A. A boolean on the ModelDeployment template, beside the existing privileged. The user names
    the privilege explicitly, which matches the member-side rule. Cost: it splits the fabric contract
    in two — the backend says EFA, the user must know to flip this on every engine consuming it, and
    nothing refuses the mismatch.
  • B. Derive it from the referenced backend. The resolve webhook already fetches the
    KVCacheBackend at mutate time, so a ModelDeployment binding a fabric backend could have the same
    access rendered into its engine container. One declaration, both sides consistent. Cost: the
    privilege becomes inferred from a binding rather than requested on the workload — the rule turns
    into "binding a fabric backend is the request", which needs an admission-time guard to stay
    honest. It also cannot stand alone for an engine consuming a remote store with no local binding.
  • C. A pool-level opt-in. The pool that admits fabric members marks itself fabric-capable and
    engines admitted against it inherit the access. Cost: the placement contract (per-role
    instanceType) becomes a security contract, which it was never meant to be.

A fourth answer is legitimate and is not the absence of a decision: members only, engines stay on
TCP. If that is the answer it has to be written as a rule, with the consequence stated — see the
sibling issue on pool-side reuse, which is what that rule costs.

Note that the field shape #286 assumed for option B no longer exists: the member's fabric device
resource is no longer a declared field but is derived from the transport protocol, so "read
deviceResourceName the same way the member rendering does" now means "derive it the same way".

What does NOT close this

  • Rendering the access without deciding whether. The ablation says what it costs, not that it
    should be granted.
  • Documenting the gap. It is documented on both pages a reader would look for it. A deployment
    that silently serves at TCP speed is not helped by a page the operator's author read.
  • A green run on a cluster where the engines do get the access. What has judgement power is what
    happens to a deployment whose engines do not — today that is a healthy-looking deployment, which
    is the thing to fix or to declare intended.
  • Choosing a shape without saying what refuses a mismatch. All three shapes admit a state where
    the backend declares a fabric and an engine does not carry the access. Whichever is chosen has to
    say whether that state is refused, reported, or allowed.

/kind enhancement
/area kv-cache

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions