Skip to content

fix(supervisor): wait for repair when the gateway refuses a startup policy write - #3785

Open
shiju-nv wants to merge 5 commits into
NVIDIA:mainfrom
shiju-nv:fix/startup-policy-write-rejection
Open

shiju-nv wants to merge 5 commits into
NVIDIA:mainfrom
shiju-nv:fix/startup-policy-write-rejection

Conversation

@shiju-nv

@shiju-nv shiju-nv commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

When the gateway refuses an image-policy upload or baseline-path write-back with FAILED_PRECONDITION or INVALID_ARGUMENT, keep the workload stopped while the user repairs its policy or provider within the provisioning deadline. Report rejection of the current configuration and continue polling after acknowledgment. If a repair makes the rejection report stale, immediately fetch the newer configuration instead of retrying the obsolete report until startup fails.

Related Issue

Fixes #3784

Changes

  • Keep the workload in Provisioning with ConfigurationInvalid after an acknowledged rejection, then read a fresh snapshot every two seconds. The gateway controls the public status message; the supervisor log records the specific write refusal.
  • Treat ABORTED from a rejection report as a stale generation and reconcile immediately. Preserve authentication failures, transient retries, and the provisioning deadline.
  • Bound and sanitize gateway messages and suppress repeated logs for the same write, gRPC code, and snapshot.
  • Apply global-policy baseline paths locally without creating a sandbox policy revision, and update policy documentation.

Testing

  • mise run pre-commit passes
  • Unit tests added/updated
  • E2E tests added/updated (if applicable)

Local verification passed package formatting, changed-file license checks, git diff --check, cargo check -p openshell-supervisor --locked, and nineteen exact library regressions. The race regressions repair desired configuration between the refused write and its rejection report for both startup write paths. They require immediate reconciliation and acceptance of the repaired generation. The retained reproductions failed on the previous head. A test-only correction boxed the two race-test futures to satisfy Clippy; both exact tests and scoped supervisor Clippy passed afterward.

Hosted Branch Checks and standard runtime E2E passed on 152c4dd9392093f3d3b0266e325e403dab6f55f1. The optional GPU and Kubernetes HA/credential-driver suites were skipped.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture docs updated (if applicable)

…olicy write

Startup writes the sandbox policy to the gateway in two cases: it
uploads a discovered image policy when the gateway has none, and it
writes the policy back after adding the proxy baseline filesystem paths.
When the gateway refused either write with FAILED_PRECONDITION or
INVALID_ARGUMENT, for example because the policy binds a provider that
is not attached, startup treated the refusal as a permanent error and
the supervisor exited. The sandbox never reached the ConfigurationInvalid
repair state that other startup rejections use.

Report such a refusal as a configuration rejection carrying the
gateway's message, log it once per write and error code, and keep
polling, so attaching the provider or replacing the policy completes
startup. Other error codes keep their current handling: transient codes
are retried, and permission, not-found and authentication failures
still end startup.

Skip the baseline-path write-back while a global policy is active. The
gateway refuses every sandbox policy write in that state, so startup
exited whenever a global policy lacked a baseline path. The supervisor
now adds the paths to its own copy of the policy without saving a
revision.

Signed-off-by: Shiju <shiju@nvidia.com>
Keep a second tracing dispatcher alive while capturing startup refusal
logs. With only one dispatcher, a parallel test thread without a default
subscriber can cache Interest::never for the shared OCSF callsite after
the capture thread rebuilds the cache.

Preserve the exact log-count, diagnostic, configuration-generation and
repair assertions. Production startup behavior is unchanged.

Signed-off-by: Shiju <shiju@nvidia.com>
Preserve upstream architecture documentation removal and retain rejected-write recovery guidance in the policy troubleshooting page.

Signed-off-by: Shiju <shiju@nvidia.com>

@johntmyers johntmyers left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gator-agent

PR Review Status

The focused supervisor repair is project-valid and its user-facing policy documentation is updated. One startup race still defeats the repair window: if an operator repairs the configuration before the rejection report lands, the gateway rejects that now-stale report and the supervisor eventually exits instead of reading the repaired generation.

Action required: @shiju-nv, handle a stale rejection report by returning to reconciliation and add coverage for repairs that race with rejection reporting.

Blocking findings:

  • GATOR-1f00b024-01: A stale rejection report can terminate startup after the operator has already repaired the configuration.

Carried findings:

  • None

Non-blocking suggestions:

  • Align tests and user-facing claims with the generic configuration status the production gateway stores for a refused write, or preserve the bounded gateway diagnostic through a trusted production path.
Gator metadata
  • Validation: Focused fix for linked issue #3784, with a reproducible supervisor startup failure and explicit acceptance criteria.
  • Docs: Relevant Fern policy guidance is updated; navigation changes are not needed.
  • Checks: Current branch, Helm, and Trivy gates are green; full E2E dispatch waits until review feedback is resolved.
  • E2E: test:e2e is required for supervisor and gateway startup interaction but is not yet dispatched while a blocker remains.
  • Head SHA: 1f00b024df7e9c4cd3aaf249430f98e57ccd6936
  • Base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Merge base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Patch ID: 2eb4d096db5a696a25aca0c05f4977e7b83dc55d
  • Gator payload: 9
  • Review mode: initial
  • Previous reviewed SHA: none
  • Review budget exhausted: no
  • Maintainer decision required: no
  • Next state: gator:in-review

Comment thread crates/openshell-supervisor/src/lib.rs
@johntmyers johntmyers added the gator:in-review Gator is reviewing or awaiting PR review feedback label Sep 29, 2026
Refetch desired configuration immediately when a rejection report is aborted because its generation changed. Preserve acknowledged rejection pacing and all other report error handling.

Signed-off-by: Shiju <shiju@nvidia.com>
@johntmyers johntmyers added the test:e2e Requires end-to-end coverage label Sep 30, 2026
@github-actions

Copy link
Copy Markdown

Label test:e2e applied for 052527f. Open the existing run and click Re-run all jobs to execute with the label set. The run will execute the standard E2E suite after building the required gateway, sandbox, and supervisor images once. The matching required CI gate status on this PR will flip green automatically once the run finishes.

@johntmyers johntmyers left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gator-agent

PR Review Status

Thanks @shiju-nv. I checked the follow-up delta against the stale rejection-report race from the previous review. The new ABORTED handling returns immediately to reconciliation, and the added regressions cover repairs racing with both refused image-policy uploads and refused baseline-path write-backs. The prior finding is resolved, with no new blocking findings in this delta.

Blocking findings:

  • No blocking findings remain

Carried findings:

  • GATOR-1f00b024-01: resolved by the current-head stale-report reconciliation and race coverage
Gator metadata
  • Validation: Focused fix for linked issue #3784, with a reproducible supervisor startup failure and explicit acceptance criteria.
  • Docs: Relevant Fern policy guidance is updated; navigation changes are not needed.
  • Checks: Current-head branch checks are running; Helm and Trivy gates are green.
  • E2E: test:e2e is applied, and the current-head Branch E2E Checks run was rerun and is queued.
  • Head SHA: 052527fc02ee71b180be88de01827f6d679afe91
  • Base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Merge base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Patch ID: 936c1e50a70dd453e80af5b3e3068c702cb18dc5
  • Gator payload: 9
  • Review mode: follow_up
  • Previous reviewed SHA: 1f00b024df7e9c4cd3aaf249430f98e57ccd6936
  • Review budget exhausted: no
  • Maintainer decision required: no
  • Next state: gator:watch-pipeline

@johntmyers johntmyers added gator:watch-pipeline Gator is monitoring PR CI/CD status gator:in-review Gator is reviewing or awaiting PR review feedback and removed gator:in-review Gator is reviewing or awaiting PR review feedback gator:watch-pipeline Gator is monitoring PR CI/CD status labels Sep 30, 2026
Keep the repair regressions below the large-future lint threshold without changing their inputs, scheduling, or assertions.

Signed-off-by: Shiju <shiju@nvidia.com>

@johntmyers johntmyers left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

gator-agent

PR Review Status

Thanks @shiju-nv. I checked the latest follow-up delta: boxing the two startup-race test futures preserves their inputs, scheduling, and assertions, and it does not reopen the resolved stale-report race. No blocking findings remain.

Blocking findings:

  • No blocking findings remain

Carried findings:

  • GATOR-1f00b024-01: remains resolved; this test-only delta does not invalidate the stale-report reconciliation or its regression coverage
Gator metadata
  • Validation: Focused fix for linked issue #3784, with a reproducible supervisor startup failure and explicit acceptance criteria.
  • Docs: Relevant Fern policy guidance is updated; navigation changes are not needed.
  • Checks: Current-head Branch Checks, Helm, and Trivy gates are green; the required E2E workflow is running.
  • E2E: test:e2e is applied, and current-head Branch E2E Checks run 36688334257 is in progress.
  • Head SHA: 152c4dd9392093f3d3b0266e325e403dab6f55f1
  • Base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Merge base SHA: 0ea0d3102089ebaa4e891a12389056f545b48415
  • Patch ID: 666233ec552d839d983a360b4a553f0e564c6156
  • Gator payload: 9
  • Review mode: follow_up
  • Previous reviewed SHA: 052527fc02ee71b180be88de01827f6d679afe91
  • Review budget exhausted: no
  • Maintainer decision required: no
  • Next state: gator:watch-pipeline

@johntmyers johntmyers added gator:watch-pipeline Gator is monitoring PR CI/CD status gator:approval-needed Gator completed review; maintainer approval needed and removed gator:in-review Gator is reviewing or awaiting PR review feedback gator:watch-pipeline Gator is monitoring PR CI/CD status labels Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gator:approval-needed Gator completed review; maintainer approval needed test:e2e Requires end-to-end coverage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Sandbox startup exits instead of waiting for repair when the gateway refuses its policy

2 participants