Follow-up to #708, which I filed in Nov 2023 and which was resolved by adding lease-strategy / LeaseStrategy.retainOnError(). Having measured the behaviour, I do not think retain-on-error delivers recovery for typical configurations, because of how the retry delay is computed.
Behaviour
When a renewal fails and the strategy retains the lease, the retry is rescheduled using the lease duration Vault last advertised, not the lease's remaining validity. The renewal interval is:
interval = max(min-renewal, advertisedDuration - expiry-threshold)
and it is recomputed after every attempt, successful or not, from the same advertised duration. Nothing accounts for the lease time already spent.
Since renewal fires when expiry-threshold of the lease remains, a failure leaves exactly expiry-threshold of runway, but the retry is scheduled a full interval later. Recovery is therefore only possible when:
2 * interval + read-timeout < ttl
which, for expiry-threshold < ttl, means:
expiry-threshold > (ttl + read-timeout) / 2
expiry-threshold reads as a safety margin, so the natural instinct is to set it small relative to the TTL. A small threshold produces a long interval, which guarantees the retry lands after the credential is dead. The configuration that looks safest is the one where retain-on-error cannot work, and it fails silently.
Retrying does continue past the first failure, but only while Vault still knows the lease. Once the credential has expired, the next attempt gets 400 lease not found, and that is terminal: the container publishes SecretLeaseExpiredEvent and stops. In a longer outage I measured renewal attempts at 14857 ms, 30881 ms and 46889 ms, the first two timing out and the last one landing after Vault was reachable again and getting the 400. Nothing subsequently re-acquires the secret, since the database backend is registered in RENEW mode, so the application cannot recover without a restart.
Real example from our deployment: a 24h database credential with expiry-threshold: 6h gives an 18h interval. Renewal fires with 6h left, fails on a transport timeout, and the retry is scheduled 18h out. The credential dies 6h after the failure, so the retry arrives 12h too late and gets 400 lease not found. Identical outcome to drop-on-error, which is what #708 was about.
Reproduction
Vault database backend, PostgreSQL, a proxy that can black-hole traffic to Vault only, and a probe taking a connection from the pool every second. Scaled down so a lease lives 20s. read-timeout: 3000. Times are relative to application start.
Does not recover. ttl 20s, min-renewal 2s, expiry-threshold 7s (interval 13s),
retain-on-error:
2102 ms scheduled renewal, lease duration=20s
14833 ms renewal attempt
17856 ms SecretLeaseErrorEvent VaultException: Cannot renew lease (+3.02s, read timeout)
17857 ms scheduled renewal, lease duration=20s <- full interval again
22760 ms pool probe fails, credential expired
30858 ms renewal attempt (+13.0s after the error)
30866 ms SecretLeaseExpiredEvent, 400 lease not found <- terminal
nothing further for the remaining 49s
Recovers. Same but expiry-threshold 15s (interval 5s), same fault window:
6005 ms fault injected
6923 ms renewal attempt
9950 ms SecretLeaseErrorEvent Cannot renew lease (+3.02s)
9951 ms scheduled renewal, lease duration=20s
11002 ms fault cleared
14952 ms renewal attempt (+5.0s after the error)
14960 ms AfterSecretLeaseRenewedEvent, renewal succeeded
renewals every 5.0s thereafter, zero probe failures
The only difference between the two runs is expiry-threshold, which crosses (20 + 3) / 2 = 11.5s.
Versions
Measured identically, to the millisecond, on:
spring-cloud-vault 4.3.2 with spring-vault-core 3.2.0 (Boot 3.5.11)
spring-cloud-vault 5.0.2 with spring-vault-core 4.1.0 (Boot 4.0.7)
The timelines above are from the second set. The harness as published builds it on Boot 4.1.0.
Suggestions
Any of these would resolve it, in decreasing order of preference:
- Record when the lease was acquired or last renewed, and compute the retry delay from the remaining validity rather than the advertised duration.
- On a retained error specifically, bound the retry so it cannot exceed the remaining window, for instance
min(interval, remaining / 2), optionally with backoff.
- At minimum, document the
expiry-threshold > (ttl + read-timeout) / 2 constraint. retain-on-error currently reads as "keeps retrying until it works". It does keep retrying, but a retry that arrives after the credential has expired can only get 400 lease not found, which ends it. For common values that means a single retry, scheduled well after the credential is already dead.
You can find an AI generated reproduction harness (podman-orchestrated Vault, PostgreSQL and a fault-injecting proxy, with the timelines above generated from it) at https://github.com/bendem/spring-cloud-vault-lease-renewal.
Follow-up to #708, which I filed in Nov 2023 and which was resolved by adding
lease-strategy/LeaseStrategy.retainOnError(). Having measured the behaviour, I do not thinkretain-on-errordelivers recovery for typical configurations, because of how the retry delay is computed.Behaviour
When a renewal fails and the strategy retains the lease, the retry is rescheduled using the lease duration Vault last advertised, not the lease's remaining validity. The renewal interval is:
and it is recomputed after every attempt, successful or not, from the same advertised duration. Nothing accounts for the lease time already spent.
Since renewal fires when
expiry-thresholdof the lease remains, a failure leaves exactlyexpiry-thresholdof runway, but the retry is scheduled a fullintervallater. Recovery is therefore only possible when:which, for
expiry-threshold < ttl, means:expiry-thresholdreads as a safety margin, so the natural instinct is to set it small relative to the TTL. A small threshold produces a long interval, which guarantees the retry lands after the credential is dead. The configuration that looks safest is the one whereretain-on-errorcannot work, and it fails silently.Retrying does continue past the first failure, but only while Vault still knows the lease. Once the credential has expired, the next attempt gets
400 lease not found, and that is terminal: the container publishesSecretLeaseExpiredEventand stops. In a longer outage I measured renewal attempts at 14857 ms, 30881 ms and 46889 ms, the first two timing out and the last one landing after Vault was reachable again and getting the 400. Nothing subsequently re-acquires the secret, since the database backend is registered inRENEWmode, so the application cannot recover without a restart.Real example from our deployment: a 24h database credential with
expiry-threshold: 6hgives an 18h interval. Renewal fires with 6h left, fails on a transport timeout, and the retry is scheduled 18h out. The credential dies 6h after the failure, so the retry arrives 12h too late and gets400 lease not found. Identical outcome todrop-on-error, which is what #708 was about.Reproduction
Vault database backend, PostgreSQL, a proxy that can black-hole traffic to Vault only, and a probe taking a connection from the pool every second. Scaled down so a lease lives 20s.
read-timeout: 3000. Times are relative to application start.Does not recover.
ttl 20s,min-renewal 2s,expiry-threshold 7s(interval 13s),retain-on-error:Recovers. Same but
expiry-threshold 15s(interval 5s), same fault window:The only difference between the two runs is
expiry-threshold, which crosses(20 + 3) / 2 = 11.5s.Versions
Measured identically, to the millisecond, on:
spring-cloud-vault4.3.2 withspring-vault-core3.2.0 (Boot 3.5.11)spring-cloud-vault5.0.2 withspring-vault-core4.1.0 (Boot 4.0.7)The timelines above are from the second set. The harness as published builds it on Boot 4.1.0.
Suggestions
Any of these would resolve it, in decreasing order of preference:
min(interval, remaining / 2), optionally with backoff.expiry-threshold > (ttl + read-timeout) / 2constraint.retain-on-errorcurrently reads as "keeps retrying until it works". It does keep retrying, but a retry that arrives after the credential has expired can only get400 lease not found, which ends it. For common values that means a single retry, scheduled well after the credential is already dead.You can find an AI generated reproduction harness (podman-orchestrated Vault, PostgreSQL and a fault-injecting proxy, with the timelines above generated from it) at https://github.com/bendem/spring-cloud-vault-lease-renewal.