Skip to content

perf(controller): react to node and endpoint changes instead of polling - #49

Open
lexfrei wants to merge 1 commit into
fix/hcloud-rate-limitfrom
perf/watch-cluster-changes
Open

lexfrei wants to merge 1 commit into
fix/hcloud-rate-limitfrom
perf/watch-cluster-changes

Conversation

@lexfrei

@lexfrei lexfrei commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Target changes now reach the balancer through watches instead of a 30-second requeue, so a balancer that nothing changed costs one or two Hetzner requests every five minutes.

A node change reconciles every service when something target selection reads has changed: labels, cordoning, readiness or addresses. Heartbeats are ignored. For robotlb balancers with the Local traffic policy, an EndpointSlice change reconciles its service when the nodes its targets follow changed. With a selector these are all endpoint nodes, since targets come from pods. Without one they are the ready ones. Readiness flaps and address changes that leave those nodes alone cost nothing.

The controller's watch mapper cannot tell a deleted slice from an updated one. A second watch, metadata only, tracks which slices exist, and a deleted slice with endpoints reconciles every service. Controller::reconcile_on could target only its service, but it is still behind an unstable kube feature, so a comment marks the place to switch.

The resync is ROBOTLB_RESYNC_INTERVAL, 300 seconds by default and between one second and a year. Finished pods and pods without an IP no longer count as targets. Neither is in the endpoint slice, so their removal would otherwise wait for the resync.

When Hetzner refuses a target for a reason other than the rate limit, the service is checked again within 30 seconds instead of at the resync, as before. A node refused for good, such as one outside the vSwitch subnet, keeps its service on that cycle. The README says so, and #54 is about telling temporary refusals from permanent ones.

Known limits: every node change that matters reconciles all services, so a node going through cordon, drain and back triggers a few rounds. Both EndpointSlice watches also run with ROBOTLB_DYNAMIC_NODE_SELECTOR=false, where they trigger nothing.

This is stacked on #37 and targets its branch.

Closes #39

Every LoadBalancer service was reconciled every 30 seconds, and each run
cost one or two Hetzner API requests even when nothing changed. The
short interval was there because only Services were watched, so a new
node or a moved pod reached the balancer only through that requeue.

Node changes that affect targets now reconcile every service. Heartbeat
updates are ignored, since only labels, cordoning, readiness and
addresses are compared. An endpoint slice change reconciles its service
when it is a robotlb balancer with the Local traffic policy and the
nodes its targets follow changed: all endpoint nodes when targets come
from pods, ready endpoint nodes otherwise. Readiness flaps and address
changes that leave those nodes alone cost nothing. The controller's
watch mapper cannot tell a deletion from an update, so a second watch
tracks which slices exist; a deleted slice with endpoints reconciles
every service. The periodic resync is now ROBOTLB_RESYNC_INTERVAL, five
minutes by default and at most a year.

A reconcile that could not add some targets, for a reason other than
the rate limit, is retried after 30 seconds like a failed one, so a
temporarily rejected node does not wait for the resync.

Pods that have finished or have no IP yet no longer count as targets.
Neither is in the endpoint slice, so deleting them would leave their
node a target until the resync.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant