Conversation
Once the Hetzner Cloud API budget ran out, every service was retried every 30 seconds, and each retry spent the budget being waited for. The budget is per project, so after a 429 all services now pause together, starting at one minute and doubling up to 16 minutes while the limit keeps being hit. Adding targets stops at the first 429 instead of calling the API for every remaining node. The generated client drops response headers, so RateLimit-Reset cannot be used. A failed reconciliation now leaves a SyncLoadBalancerFailed warning event on the service, so kubectl describe shows the last error. Errors are logged on one line with the status and the message Hetzner reported, and the API token, which Hetzner quotes in some messages, is redacted from logs and events. The chart role gains permission to create events. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
Services paused by the rate limit gate were all requeued for the moment it reopened, so they hit the API in one burst while the budget was still nearly empty. Each wait now grows by up to a quarter, derived from the service name, so services wake up spread over that window. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
With wakeups spread by service name, the same service wakes first after every pause, hits the limit and closes the gate again. Every other service only ever saw the closed gate, which published no event and logged at debug level, so after the event TTL they showed nothing at all. A service that wakes up to a closed gate now gets an event and an info line. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
The comment claimed the default grants all permissions, while it lists only what robotlb needs. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
lexfrei
force-pushed
the
fix/hcloud-rate-limit
branch
from
September 29, 2026 09:47
a1547b9 to
785e5bf
Compare
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Once the Hetzner Cloud API budget runs out, robotlb retries every service every 30 seconds, and each retry spends the budget it is waiting for. The failure is also hard to see from the cluster.
The budget is per project, so after a 429 all services now pause together. The pause starts at one minute and doubles up to 16 minutes while the limit keeps being hit. A 429 from a call that was already in flight does not lengthen it, and adding targets stops at the first 429 instead of calling the API for every remaining node. Services wake up spread over up to a quarter of the pause, so they don't all hit the API at the moment it ends.
RateLimit-Resetis not used: the generated hcloud client drops response headers, so the header never reaches robotlb.A failed reconciliation now leaves a
SyncLoadBalancerFailedwarning event on the service, the same reason the upstream service controller uses, sokubectl describeshows the last error. A service that wakes up to the paused gate gets one too. Errors are logged on one line with the HTTP status and the code and message Hetzner reported. The API token, which Hetzner quotes in some error messages, is redacted from logs and events. The chart role gets permission to create events.Charts that override
serviceAccount.permissionsreplace the default list, so they needcreateonevents.k8s.ioevents added by hand. Without it, a failure only logs that the event could not be published.Other errors are still retried every 30 seconds, see #46, and kube 0.96 does not aggregate events, see #45.
Closes #36