User Story
As an operator, I want sandbox restart and recovery to behave consistently across drivers, so delayed cleanup cannot disrupt a newer run.
Problem Statement
Lifecycle coordination is driver-specific. Background observations or cleanup must not stop, delete, or invalidate a newer runtime. Reused resource IDs do not identify a run. Protecting runtime resources and filtering stale published state are separate guarantees: gateway gates/event filters cannot undo containment already performed inside a driver.
Impact / Why This Matters
Podman reconciliation can stop a newly restarted workload while its supervisor is still exited. Retrying is unreliable if cleanup races recur; isolated fixes leave equivalent paths unaudited. The focused fix is PR #4117. Equivalent failures in other drivers remain unproven.
Proposed Design
- Land the focused Podman fix. Extract an instance-local gate independently of engine-specific exit timestamps. Specify lock lifetime, clone sharing, identity revalidation, rollback and genuine-loss containment. Reconciliation may wait (Podman) or skip a busy sandbox and retry (Kubernetes); document that choice.
- Audit/adopt one driver at a time: lifecycle RPCs, watcher/initial sync, supervisor-loss, get/list side effects, admission, cleanup, cancellation and retry. Preserve readiness, exec, workspace retention and genuine-loss containment.
- Separately design runtime-generation ownership and completion/fencing of accepted remote mutations across cancellation, independent processes and HA. A local mutex alone cannot supply these guarantees.
Acceptance Criteria
Alternatives Considered
Continue isolated fixes: small changes but inconsistent guarantees. One cross-driver refactor: combines local coordination with distributed ownership design and delays the proven Podman fix.
Implementation and Test References
Repeatable Live Qualification
Use the tmachine workflow: nix run .#tmachine -- setup fedora-podman-rootless, build current Linux artifacts with nix run .#build-artifacts-binaries, then nix run .#tmachine -- test fedora-podman-rootless none shell. Run a rootless Podman gateway with current sandbox/supervisor binaries and a workload with matching unprivileged UID/GID. Create and stop a Ready sandbox; collect labeled Podman events and both container states during restart.
In a disposable test build, pause three seconds after workload start, before supervisor start. Compare with only the start gate bypassed versus retained; remove instrumentation afterward. On Fedora 44 / Podman 5.8.1 / SELinux Enforcing, bypassing stopped the workload after 254 ms and returned ContainerExited; retaining reached Ready and successful exec. Kill the supervisor afterward and verify workload containment. Live gateway recovery after that forced loss was rejected by phase preconditions; driver-level retry tests do not establish gateway recovery.
For each other driver, place a test barrier at its equivalent partial-start/cleanup boundary, delay old reconciliation across restart, verify no premature destructive action and eventual Ready/exec, then inject genuine dependency loss. Audit every destructive actor and repeat on that driver's live runtime.
User Story
As an operator, I want sandbox restart and recovery to behave consistently across drivers, so delayed cleanup cannot disrupt a newer run.
Problem Statement
Lifecycle coordination is driver-specific. Background observations or cleanup must not stop, delete, or invalidate a newer runtime. Reused resource IDs do not identify a run. Protecting runtime resources and filtering stale published state are separate guarantees: gateway gates/event filters cannot undo containment already performed inside a driver.
Impact / Why This Matters
Podman reconciliation can stop a newly restarted workload while its supervisor is still exited. Retrying is unreliable if cleanup races recur; isolated fixes leave equivalent paths unaudited. The focused fix is PR #4117. Equivalent failures in other drivers remain unproven.
Proposed Design
Acceptance Criteria
Alternatives Considered
Continue isolated fixes: small changes but inconsistent guarantees. One cross-driver refactor: combines local coordination with distributed ownership design and delays the proven Podman fix.
Implementation and Test References
restart_and_reconciliation_do_not_stop_the_new_workload; companion tests cover cancellation/failure/retry, delayed events and admission. Runcargo test -p openshell-driver-podman.Repeatable Live Qualification
Use the tmachine workflow:
nix run .#tmachine -- setup fedora-podman-rootless, build current Linux artifacts withnix run .#build-artifacts-binaries, thennix run .#tmachine -- test fedora-podman-rootless none shell. Run a rootless Podman gateway with current sandbox/supervisor binaries and a workload with matching unprivileged UID/GID. Create and stop a Ready sandbox; collect labeled Podman events and both container states during restart.In a disposable test build, pause three seconds after workload start, before supervisor start. Compare with only the start gate bypassed versus retained; remove instrumentation afterward. On Fedora 44 / Podman 5.8.1 / SELinux Enforcing, bypassing stopped the workload after 254 ms and returned
ContainerExited; retaining reached Ready and successful exec. Kill the supervisor afterward and verify workload containment. Live gateway recovery after that forced loss was rejected by phase preconditions; driver-level retry tests do not establish gateway recovery.For each other driver, place a test barrier at its equivalent partial-start/cleanup boundary, delay old reconciliation across restart, verify no premature destructive action and eventual Ready/exec, then inject genuine dependency loss. Audit every destructive actor and repeat on that driver's live runtime.