Conversation
Every write to a Service starts a reconcile, including status writes by another LoadBalancer controller that also takes Services without a class. When that controller clears status.loadBalancer, robotlb writes it back, the other controller clears it again, and each round costs Hetzner requests until the project hits the rate limit. Services carry no metadata.generation, so a generation predicate cannot filter these events, and filtering the main stream needs an unstable kube-runtime feature. robotlb now remembers, per Service UID, a hash of the spec, the robotlb/ annotations and the computed targets and ports after each successful reconcile. A reconcile that finds the same hash before the next check is due makes no Hetzner requests and writes no status. Node and EndpointSlice changes still get through because they change the targets. An entry is valid until the requeue of the reconcile that recorded it, so while Hetzner refuses some targets events are skipped only until the 30-second retry. The entry is dropped after a failed reconcile, when the Service is released or deleted, and when it has no port to expose. The README now says which classes robotlb handles, that setting loadBalancerClass: robotlb keeps other controllers away, and that recreating a Service to add the class costs it its balancer and public IP. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
Collaborator
Author
|
Checked on a live cluster with a build of this branch and debug logging, on a LoadBalancer Service with a nodePort:
The API budget can't show the difference over a window this short, since it refills by one request per second. The skip line and the shorter log per reconcile are the evidence. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
robotlb now skips the Hetzner calls when nothing it acts on has changed since the last successful reconcile. The README now explains how to keep another LoadBalancer controller away from robotlb's Services.
Two controllers that both take Services without a class keep undoing each other's
status.loadBalancer. MetalLB started without--lb-classis one example. Every status write is a watch event and starts a reconcile, and every reconcile talks to Hetzner. The loop runs until the project hits the rate limit. Services have nometadata.generation, so a generation predicate does not filter these events. A filter on the main stream needsController::for_stream, which is still unstable.After each successful reconcile, robotlb keeps a hash per Service UID. It covers the spec, the
robotlb/annotations, and the targets and ports robotlb computed. The next reconcile computes targets and ports as before, which costs only Kubernetes reads. If the hash is the same and the next check is not due yet, it stops before the first Hetzner call and does not write the status. It requeues for the time left until that check. The next check is the resync, or the 30-second retry while Hetzner refuses some targets. Node and EndpointSlice changes still get through, because they change the targets.The hash is dropped after a failed reconcile, and when the Service is released or deleted. It is also dropped when the Service has no port to expose, because that path clears the status and the status has to come back with the ports.
The README now says that robotlb handles Services with no
loadBalancerClassor withrobotlb. SettingloadBalancerClass: robotlbkeeps the other controller away. The field can only be set when the Service is created or its type is changed toLoadBalancer. Recreating a Service deletes its balancer, and the new one gets a new public IP, so the README says to set the class at creation.There are two trade-offs:
robotlbclass ends the fight over the status.Some smaller limits I left alone. If a Service is deleted without robotlb seeing its deletion (for example after someone removes the finalizer by hand), its hash stays in memory until the process restarts. If the balancer has no public IP yet on the first reconcile, the status is written at the next check, not at the next Service event. A reconcile that waits at the rate limit gate also counts as failed and drops the hash, so every Service does a full reconcile once the gate opens. Keeping the hash there would save requests right when the budget is lowest, but I kept "any failure drops the hash" as the simpler rule.
Unit tests cover the skip decision before and after the check is due, and with the same or a different hash. They check that the hash changes with a spec field, a robotlb annotation, the targets and a port. They also check that it stays the same when only the status,
resourceVersionor a foreign annotation changes, and that the hash is dropped on failure, release and deletion. I broke the code under each test and saw it go red: without the targets, without the annotations, with the expiry ignored, with a fixed expiry instead of the one recorded, and with the hash kept after a failure. Two mutations survive. One records the hash before the Hetzner calls succeed, the other bypasses the check. Both sit inreconcile_load_balancer, and catching them needs a mocked Kubernetes API and Hetzner API. A failed reconcile drops the hash anyway, so the first one only changes how long the hash stays valid when Hetzner refuses some targets.Stacked on #56, only the last commit is new.
Closes #69