Conversation
…olicy write Startup writes the sandbox policy to the gateway in two cases: it uploads a discovered image policy when the gateway has none, and it writes the policy back after adding the proxy baseline filesystem paths. When the gateway refused either write with FAILED_PRECONDITION or INVALID_ARGUMENT, for example because the policy binds a provider that is not attached, startup treated the refusal as a permanent error and the supervisor exited. The sandbox never reached the ConfigurationInvalid repair state that other startup rejections use. Report such a refusal as a configuration rejection carrying the gateway's message, log it once per write and error code, and keep polling, so attaching the provider or replacing the policy completes startup. Other error codes keep their current handling: transient codes are retried, and permission, not-found and authentication failures still end startup. Skip the baseline-path write-back while a global policy is active. The gateway refuses every sandbox policy write in that state, so startup exited whenever a global policy lacked a baseline path. The supervisor now adds the paths to its own copy of the policy without saving a revision. Signed-off-by: Shiju <shiju@nvidia.com>
Keep a second tracing dispatcher alive while capturing startup refusal logs. With only one dispatcher, a parallel test thread without a default subscriber can cache Interest::never for the shared OCSF callsite after the capture thread rebuilds the cache. Preserve the exact log-count, diagnostic, configuration-generation and repair assertions. Production startup behavior is unchanged. Signed-off-by: Shiju <shiju@nvidia.com>
Discard snapshot-derived startup writes and rejection reports on ABORTED. Keep transient transport retries and permanent errors unchanged, and retain the reconciliation budget when a rejected report is already obsolete. Cover both write races and stale rejection reports, and verify image-policy recovery after provider attachment without restarting or relaunching work. Signed-off-by: Shiju <shiju@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fetch fresh configuration when the gateway returns
ABORTEDfor a startup policy write or rejection report. Repeating the stale request can overwrite an operator's repair or exhaust retries and terminate startup.Related Issue
Follow-up to #3785 and related to #3784. This is a localized correction to startup conflict recovery; it does not add a conditional policy-write API.
Changes
Testing
mise run pre-commitpassesChecklist