Skip to content

robotlb and another LoadBalancer controller reconcile the same Service in a loop until the Hetzner rate limit #69

Description

@lexfrei

When another LoadBalancer controller also handles a Service, robotlb and that controller keep undoing each other's status writes, and robotlb spends the whole Hetzner API budget on it. I'd like robotlb to skip the Hetzner calls when nothing it acts on has changed, and the README to explain how to keep other controllers away from robotlb's Services.

I see this next to MetalLB installed with no address pool. It happens on 0.0.5, on 0.0.6, and on a build with the open rate limit PRs. robotlb reconciles the Service and writes the balancer IP into status.loadBalancer. In the same second MetalLB finds no pool for the Service and updates it, and the IP is gone again. About 1 ms later robotlb starts the next reconcile with reason "object updated". One round takes about 200 ms. In 20 minutes one Service was reconciled 330 times and another 176 times, and 1008 of the triggers were "object updated". The loop stops only when Hetzner answers 429 (3600 requests per hour). The pause from #37 grows from 60 s to 960 s, and the loop starts again when it ends. The balancers themselves are fine and their targets are correct, but the status of all five LoadBalancer Services is empty.

The controller watches every Service with no filter, so any write to a Service starts a reconcile, a status write from another controller included (main.rs). Each reconcile talks to Hetzner before it can tell that nothing changed. The requeue after a successful reconcile does not slow this down, because a watch event starts the next reconcile right away.

Both controllers think they own these Services. robotlb takes every Service with no loadBalancerClass (main.rs). MetalLB started without --lb-class does the same (service_controller.go). Setting spec.loadBalancerClass: robotlb already works: robotlb keeps the Service and MetalLB skips it. The README does not mention this class, and I found it only in the code. The field can only be set when the Service is created or switched to LoadBalancer, so existing Services have to be recreated.

Filtering the watch by metadata.generation would not help. None of the five Services has a generation on Kubernetes 1.36, and that is by design: the API server never sets one on a Service (strategy.go, v1.37.1). So predicates::generation returns nothing and predicate_filter lets every event through. Putting a filter in front of the main stream also needs Controller::for_stream, which is still behind the unstable-runtime-stream-control feature, both in kube-runtime 0.96 and in the current release.

What I'd do instead is remember, per Service, a hash of what the reconcile acts on: the spec, the robotlb annotations, and the targets and ports it computed. Status, resourceVersion and managedFields stay out of it, so a write by the other controller does not change the hash. If the hash matches the last successful reconcile and the resync interval has not passed, the Hetzner calls are skipped. Node and EndpointSlice changes from #49 still get through, since they change the targets. The extra reconcile that robotlb triggers with its own status write then costs no API calls either.

This stops the quota burn but not the fight over the status. After the other controller clears it, the status stays empty until the next resync, and gets cleared again after it. Only a separate class fixes that, so I think the README should say that robotlb takes Services with no class or the robotlb class, and that the class is required when another LoadBalancer controller runs in the cluster.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions