Conversation
This was referenced Sep 29, 2026
lexfrei
force-pushed
the
perf/watch-cluster-changes
branch
2 times, most recently
from
September 29, 2026 15:28
26996c9 to
c4b6cd9
Compare
Every LoadBalancer service was reconciled every 30 seconds, and each run cost one or two Hetzner API requests even when nothing changed. The short interval was there because only Services were watched, so a new node or a moved pod reached the balancer only through that requeue. Node changes that affect targets now reconcile every service. Heartbeat updates are ignored, since only labels, cordoning, readiness and addresses are compared. An endpoint slice change reconciles its service when it is a robotlb balancer with the Local traffic policy and the nodes its targets follow changed: all endpoint nodes when targets come from pods, ready endpoint nodes otherwise. Readiness flaps and address changes that leave those nodes alone cost nothing. The controller's watch mapper cannot tell a deletion from an update, so a second watch tracks which slices exist; a deleted slice with endpoints reconciles every service. The periodic resync is now ROBOTLB_RESYNC_INTERVAL, five minutes by default and at most a year. A reconcile that could not add some targets, for a reason other than the rate limit, is retried after 30 seconds like a failed one, so a temporarily rejected node does not wait for the resync. Pods that have finished or have no IP yet no longer count as targets. Neither is in the endpoint slice, so deleting them would leave their node a target until the resync. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
lexfrei
force-pushed
the
perf/watch-cluster-changes
branch
from
September 29, 2026 15:32
c4b6cd9 to
dc37dcf
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Target changes now reach the balancer through watches instead of a 30-second requeue, so a balancer that nothing changed costs one or two Hetzner requests every five minutes.
A node change reconciles every service when something target selection reads has changed: labels, cordoning, readiness or addresses. Heartbeats are ignored. For robotlb balancers with the Local traffic policy, an EndpointSlice change reconciles its service when the nodes its targets follow changed. With a selector these are all endpoint nodes, since targets come from pods. Without one they are the ready ones. Readiness flaps and address changes that leave those nodes alone cost nothing.
The controller's watch mapper cannot tell a deleted slice from an updated one. A second watch, metadata only, tracks which slices exist, and a deleted slice with endpoints reconciles every service.
Controller::reconcile_oncould target only its service, but it is still behind an unstable kube feature, so a comment marks the place to switch.The resync is
ROBOTLB_RESYNC_INTERVAL, 300 seconds by default and between one second and a year. Finished pods and pods without an IP no longer count as targets. Neither is in the endpoint slice, so their removal would otherwise wait for the resync.When Hetzner refuses a target for a reason other than the rate limit, the service is checked again within 30 seconds instead of at the resync, as before. A node refused for good, such as one outside the vSwitch subnet, keeps its service on that cycle. The README says so, and #54 is about telling temporary refusals from permanent ones.
Known limits: every node change that matters reconciles all services, so a node going through cordon, drain and back triggers a few rounds. Both EndpointSlice watches also run with
ROBOTLB_DYNAMIC_NODE_SELECTOR=false, where they trigger nothing.This is stacked on #37 and targets its branch.
Closes #39