Skip to content

fix(auth): the JWKS first-fetch retry has no jitter, so every tenant retries in lockstep #674

Description

@EricAndrechek

Area: auth · footgun · found via #616's "Left for later" section

Expected: a retry loop against a shared external dependency spreads its attempts, the way internal/discovery has since #616.

Actual: the JWKS first-fetch retry is a bare doubling schedule with no jitter — internal/auth/auth.go:179 starts at jwksRetryMin (1 s), waits on <-time.After(backoff) at :195, and does backoff = min(backoff*2, jwksRetryMax) at :197, capping at 1 minute. No rand, unlike internal/discovery/discovery.go:242-243,462 ("full jitter"). Since #604 there is one verifier per tenant, each started on its own goroutine (auth.go:151), so every tenant configured against the same jwks_url begins its loop when the process boots.

Impact: while an identity provider is unreachable, N tenants sharing a jwks_url converge on the 60 s cap and send a synchronized burst of N requests every minute, instead of a spread load — the herd internal/discovery was given jitter to avoid (#141, #616). Every verifier stays pending throughout, so tokens against them get 503 + Retry-After: 30.

Scope: #616's own "Left for later" names the direction — "internal/auth (the JWKS retry) has its own unjittered loop … Consolidating them is a follow-up." discovery.go's jitter helper is the thing to reuse.

Related: #616, #141, #604, #228


From the merged-PR follow-up sweep (#616 "Left for later"); validated by code-read against 4ff50745 on 2026-09-28.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/authAuthentication: tokens, JWT/JWKS, keys, token expiry/revocationbugSomething isn't working

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions