Skip to content

enhancement: stop admitted workloads from tolerating every node taint #599

Description

@thxCode

What would you like to be added:

Every ResourceFlavor the operator derives carries tolerations: [{operator: Exists}]
(pkg/worker/controllers/worker/node_flavor.go, the flavor spec; the rationale in the comment and in
docs/architecture/scheduling-chain.md, "eligibility by nodeLabels, not taints"). Kueue copies a
flavor's tolerations into every Pod it admits through that flavor, so each admitted workload tolerates
every node taint.

Replace the blanket toleration with one that still lets quota routing ignore the taints the operator
does not care about, but stops tolerating the taints that mean "do not run here now":

  • node.kubernetes.io/unschedulable (cordon, drain, autoscaler scale-down)
  • node.kubernetes.io/not-ready and node.kubernetes.io/unreachable with NoExecute, so the default
    300 s eviction applies again
  • the cluster-autoscaler and Karpenter disruption taints (ToBeDeletedByClusterAutoscaler,
    karpenter.sh/disrupted)

The shape of the replacement (an explicit allow-list of control-plane and accelerator taints, a
setting, or keeping Exists but excluding the lifecycle taints in the Pod webhook) is open.

Why is this needed:

  • Measured: on a local kind cluster whose control-plane node carries NoSchedule, a Workload admitted
    through a derived queue was placed by topology-aware scheduling on that control-plane node and
    reserved quota there. The blanket toleration makes TAS treat every taint as tolerated.
  • Read (not yet measured on a real accelerator node): because the Pod carries an empty-key Exists
    toleration, the DefaultTolerationSeconds admission plugin no longer adds the 300 s not-ready /
    unreachable NoExecute tolerations. A serving Pod on a node that goes unreachable is therefore
    never evicted and stays bound to the dead node until its Node object is deleted.
  • Inferred: a serving Pod can be placed on a node that is cordoned for maintenance or being drained,
    and a node being scaled in by an autoscaler keeps receiving new replicas.
  • The same failure shape was fixed for the chart's own Deployments (Kueue controller manager, NFD
    master, NFD gc) in fix(chart): stop Kueue and NFD from tolerating every taint #586; this is the workload side of it.

Where the need came from (tick one):

  • Blocked — something cannot be done today; name it, and who ran into it
  • Asked for — somebody outside this repository asked for it
  • Anticipated — nobody is blocked yet; this is expected to be needed

Found during the operator's own real-cluster and kind verification rounds, not reported by a user.
The people affected are cluster administrators who cordon, drain or autoscale accelerator nodes,
and anyone relying on Kubernetes' default eviction from failed nodes.

Completion requirements:

This enhancement requires the following artifacts:

  • Design doc
  • API change
  • Docs update

The artifacts should be linked in subsequent comments.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions