You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Every ResourceFlavor the operator derives carries tolerations: [{operator: Exists}]
(pkg/worker/controllers/worker/node_flavor.go, the flavor spec; the rationale in the comment and in docs/architecture/scheduling-chain.md, "eligibility by nodeLabels, not taints"). Kueue copies a
flavor's tolerations into every Pod it admits through that flavor, so each admitted workload tolerates
every node taint.
Replace the blanket toleration with one that still lets quota routing ignore the taints the operator
does not care about, but stops tolerating the taints that mean "do not run here now":
node.kubernetes.io/not-ready and node.kubernetes.io/unreachable with NoExecute, so the default
300 s eviction applies again
the cluster-autoscaler and Karpenter disruption taints (ToBeDeletedByClusterAutoscaler, karpenter.sh/disrupted)
The shape of the replacement (an explicit allow-list of control-plane and accelerator taints, a
setting, or keeping Exists but excluding the lifecycle taints in the Pod webhook) is open.
Why is this needed:
Measured: on a local kind cluster whose control-plane node carries NoSchedule, a Workload admitted
through a derived queue was placed by topology-aware scheduling on that control-plane node and
reserved quota there. The blanket toleration makes TAS treat every taint as tolerated.
Read (not yet measured on a real accelerator node): because the Pod carries an empty-key Exists
toleration, the DefaultTolerationSeconds admission plugin no longer adds the 300 s not-ready / unreachableNoExecute tolerations. A serving Pod on a node that goes unreachable is therefore
never evicted and stays bound to the dead node until its Node object is deleted.
Inferred: a serving Pod can be placed on a node that is cordoned for maintenance or being drained,
and a node being scaled in by an autoscaler keeps receiving new replicas.
Blocked — something cannot be done today; name it, and who ran into it
Asked for — somebody outside this repository asked for it
Anticipated — nobody is blocked yet; this is expected to be needed
Found during the operator's own real-cluster and kind verification rounds, not reported by a user.
The people affected are cluster administrators who cordon, drain or autoscale accelerator nodes,
and anyone relying on Kubernetes' default eviction from failed nodes.
Completion requirements:
This enhancement requires the following artifacts:
Design doc
API change
Docs update
The artifacts should be linked in subsequent comments.
What would you like to be added:
Every
ResourceFlavorthe operator derives carriestolerations: [{operator: Exists}](
pkg/worker/controllers/worker/node_flavor.go, the flavor spec; the rationale in the comment and indocs/architecture/scheduling-chain.md, "eligibility by nodeLabels, not taints"). Kueue copies aflavor's tolerations into every Pod it admits through that flavor, so each admitted workload tolerates
every node taint.
Replace the blanket toleration with one that still lets quota routing ignore the taints the operator
does not care about, but stops tolerating the taints that mean "do not run here now":
node.kubernetes.io/unschedulable(cordon, drain, autoscaler scale-down)node.kubernetes.io/not-readyandnode.kubernetes.io/unreachablewithNoExecute, so the default300 s eviction applies again
ToBeDeletedByClusterAutoscaler,karpenter.sh/disrupted)The shape of the replacement (an explicit allow-list of control-plane and accelerator taints, a
setting, or keeping
Existsbut excluding the lifecycle taints in the Pod webhook) is open.Why is this needed:
NoSchedule, a Workload admittedthrough a derived queue was placed by topology-aware scheduling on that control-plane node and
reserved quota there. The blanket toleration makes TAS treat every taint as tolerated.
Existstoleration, the
DefaultTolerationSecondsadmission plugin no longer adds the 300 snot-ready/unreachableNoExecutetolerations. A serving Pod on a node that goes unreachable is thereforenever evicted and stays bound to the dead node until its
Nodeobject is deleted.and a node being scaled in by an autoscaler keeps receiving new replicas.
master, NFD gc) in fix(chart): stop Kueue and NFD from tolerating every taint #586; this is the workload side of it.
Where the need came from (tick one):
Found during the operator's own real-cluster and kind verification rounds, not reported by a user.
The people affected are cluster administrators who cordon, drain or autoscale accelerator nodes,
and anyone relying on Kubernetes' default eviction from failed nodes.
Completion requirements:
This enhancement requires the following artifacts:
The artifacts should be linked in subsequent comments.