Skip to content

Keep redialing a restarted codespace and bring back all its loops - #616

Merged
scgopi merged 2 commits into
mainfrom
fix/remote-reboot-recovery
Oct 3, 2026
Merged

scgopi merged 2 commits into
mainfrom
fix/remote-reboot-recovery

Conversation

@scgopi

@scgopi scgopi commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Summary

Loops on a codespace that restarted no longer stay dead until a human selects one. A down codespace is still dialed once a minute after the 4-minute pause. When it answers again, every loop on it comes back, including finished loops and the workers inside a running composite. Plain ssh hosts are swept every 30 seconds, and their finished loops come back with no pane open.

Why

After a codespace restart, loops did not recover on their own. Someone had to create or select a loop to bring them back. Reading the restore path found three gaps:

Gap Effect
CodespaceDialBreaker returned .paused forever after 240s of failed dials, and only the reconnect marker cleared it A restart slower than about 4 minutes stopped every ensure and the 60s liveness sweep for good. Selecting a loop touches the marker, which is why creating a loop "fixed" it.
A finished loop was restored only after a pane redialed (RebootProbeGate) With no pane open, finished loops never came back, on codespaces and plain ssh hosts alike.
ensureUnattendedSessionsAlive swept only graph.nodes Workers of a piloted or armed composite live on its sub-graph and were never re-ensured after a reboot.

Changes

  • CodespaceDialSchedule.slowRetryInterval (60s). Paused, the breaker lets through one dial per interval for the whole codespace, whichever read or ensure asks first. A failed slow dial leaves the outage clock alone. Enter or selecting a loop still reconnects at once.
  • Recovery hook. When a codespace answers after an outage, the breaker touches the host's redial stamp (ZmxSessionLauncher.markRedialed). The next sweep then runs the reboot probe and restores finished loops, as if a pane had redialed.
  • Pane. The paused banner now says it retries every 60s and redials on its own after that interval, on the same outage clock.
  • Sweep. It also ensures the unresolved unattended children of a piloted or armed composite, the same children pilotComposite starts.
  • Plain ssh hosts. The liveness sweep runs every 30s (was 60s). Codespaces are swept on every second tick, so they stay at 60s. RebootProbeGate always lets a plain ssh host through, so its finished loops are probed on every sweep. Those dials are multiplexed over the ControlMaster and spend no quota.

Cost while a codespace stays down past 4 minutes: about one gh dial per minute from the daemon, plus one per open pane of that codespace. A plain ssh host with finished loops costs one multiplexed probe dial every 30s. A dial to a stopped codespace starts it, so a codespace with running loops is brought back up rather than left stopped.

Tests

RED: xcodebuild test -only-testing CodespaceDialScheduleTests CodespaceSelectionTests RemoteSessionResumeTests RemoteRebootRestoreTests with SSHReconnectLoop.swift and GraphStore.swift from origin/main and the breaker's slow retry and recovery hook removed -> 4 failed of 52 (aPausedCodespaceIsStillDialedOncePerSlowInterval, aCodespaceThatAnswersAfterAnOutageAsksForTheRebootProbe, aPausedPaneRedialsOnItsOwnAfterTheSlowInterval, theLivenessSweepRestartsTheLoopsInsideARunningComposite), TEST FAILED exit 65
GREEN: same focused xcodebuild test on this branch -> 52 tests in 4 suites passed, exit 0
GREEN: after the interval change, the same focused suites -> 54 tests in 4 suites passed, exit 0; its new tests (aPlainSSHHostIsProbedOnEverySweepWithNoPaneOpen, plainSSHHostsAreSweptEveryThirtySecondsAndCodespacesEveryMinute) fail to compile on 91a930c -> ProjectRegistry has no member sweeps
REGRESSION: full xcodebuild test of the graphcode scheme -> 2226 tests in 255 suites, 1 failure (MessageDeliveryTests.aSessionWhoseTaskEndedIsNeverTypedInto), which fails identically 3 of 3 runs on base c1febe4 without this change. graphcoded and graphcode-cli schemes build, make check exit 0.

The MessageDeliveryTests failure is pre-existing on main and unrelated: it is a live-zmx send test, and this PR does not touch the send path.

scgopi and others added 2 commits October 3, 2026 08:52
A codespace that stayed unreachable for four minutes paused every dial
until a human selected one of its loops, so a restart that took longer
left every loop on it dead. The pause now lets one dial through every
five minutes, in the daemon and in an open pane. The dial that reaches
the codespace again asks for the reboot probe, so finished loops are
restored even with no pane open, and the liveness sweep now also
re-ensures the workers of a piloted or armed composite.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
The paused codespace now redials once a minute instead of every five.
Plain ssh hosts are swept every 30 seconds, and their finished loops are
probed on every sweep rather than only after a pane redials: their dials
ride the ControlMaster and spend no quota. Codespaces stay on a minute.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: scgopi <scgopireddy@gmail.com>
@scgopi
scgopi merged commit bc1d1ed into main Oct 3, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant