Skip to content

fix(install): stop waiting on a gateway unit that has failed - #4092

Closed
Yi-111-a wants to merge 1 commit into
NVIDIA:mainfrom
Yi-111-a:fix/installer-stop-waiting-on-failed-gateway-unit
Closed

Yi-111-a wants to merge 1 commit into
NVIDIA:mainfrom
Yi-111-a:fix/installer-stop-waiting-on-failed-gateway-unit

Conversation

@Yi-111-a

@Yi-111-a Yi-111-a commented Oct 2, 2026

Copy link
Copy Markdown

Summary

The installer now stops waiting as soon as the openshell-gateway user service
fails during startup, instead of burning its full 30s listener timeout on a
gateway that can never start.

Related Issue

Closes #4040

Changes

  • install.sh: added gateway_user_service_field and gateway_user_service_failed.
    The gateway is a systemd user service with Restart=on-failure and
    RestartSec=5s, so a gateway that fails during startup never settles in a
    final failed state: systemd keeps cycling it through auto-restart, and a 5s
    delay never trips the start limit that would eventually leave it failed. The
    probe therefore only saw a listener that never opens.

  • start_user_gateway now records NRestarts before restarting the unit.
    NRestarts is never reset by a restart, so reading it first is what lets the
    check tell a failure from this install apart from one left over from an
    earlier run.

  • wait_for_local_gateway_listener checks the unit on every pass of the wait
    loop. A final failed unit, and an auto-restart unit whose NRestarts has
    grown past that baseline, both end the wait early. The check runs on every
    pass, including the passes where the mTLS bundle check short-circuits.

  • The final line now names the service and gives the command to retry it, after
    dump_local_gateway_diagnostics has already printed the gateway's own error:

    openshell: error: the openshell-gateway service failed to start; retry it with: systemctl --user restart openshell-gateway
    

Per the issue, this only stops the installer from waiting. It does not stop the
unit's restart loop and does not pre-check for Docker or Podman.

macOS and snap installs are unchanged: the check is limited to Linux non-snap
installs, which is the only path that runs the gateway as a systemd user service.
macOS goes through the same wait_for_local_gateway_listener but is a
Homebrew/launchd service, and snap has its own
wait_for_snap_gateway_listener.

Testing

  • Checks appropriate to the affected code and behavior pass
  • Unit tests added/updated (if applicable)
  • E2E tests added/updated (if applicable)

mise run test:install-sh covers all four unit shapes the issue asks for —
failed, restarting after a failure, starting, and running — plus a restart loop
that predates the installer's own restart, a missing baseline, a non-numeric
restart count, macOS, and snap.

The regression is covered directly: with the in-loop check removed the probe
waits the full 30s, and the new timing assertion fails with
the listener probe waited 30s on a failed gateway unit. With the fix the same
case completes in under 5s.

Verified locally:

$ sh -n install.sh                      # POSIX sh syntax
$ bash tasks/scripts/test-install-sh.sh
install.sh focused tests passed
$ shellcheck -s sh install.sh

install.sh reports the same two pre-existing SC2016 notes before and after
this change, and no new findings. test-install-sh.sh gains only the
SC2034/SC2329 notes already present in that file for variables and override
functions consumed by the sourced install.sh (the same notes as
OPENSHELL_SNAP_TLS_DIR at line 496 and UPGRADE_NOTICE_ACK at line 201).

The gateway runs as a systemd user service with Restart=on-failure and
RestartSec=5s, so a gateway that fails during startup never settles in a
final "failed" state: systemd keeps cycling it through "auto-restart",
and a 5s delay never trips the start limit that would eventually leave it
failed. The listener probe therefore burned its whole 30s timeout on a
gateway that could never start, and its last line reported a bare
listener timeout even though the gateway's own error was in the
diagnostics printed just above it.

Probe the unit on every pass of the wait loop. A final "failed" unit, and
an "auto-restart" unit whose NRestarts has grown past the count taken
before the installer restarted it, both end the wait early. The
comparison is against that baseline so a unit that only failed during an
earlier run is still waited on, and starting or running units are never
treated as failures.

macOS and snap installs keep their existing behaviour: the check is
limited to Linux non-snap installs, which is the only path that runs the
gateway as a systemd user service.

On failure the last line now names the service and gives the command to
retry it, after the diagnostics have already shown the gateway's error.

Closes NVIDIA#4040

Signed-off-by: Yi-111-a <153097222+Yi-111-a@users.noreply.github.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Thank you for your interest in contributing to OpenShell, @Yi-111-a.

This project uses a vouch system for first-time contributors. Before submitting a pull request, you need to be vouched by a maintainer.

To get vouched:

  1. Open a Vouch Request discussion.
  2. Describe what you want to change and why.
  3. Write in your own words — do not have an AI generate the request.
  4. A maintainer will comment /vouch if approved.
  5. Once vouched, open a new PR (preferred) or reopen this one after a few minutes.

See CONTRIBUTING.md for details.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Thank you for your submission! We ask that you sign our Developer Certificate of Origin before we can accept your contribution. You can sign the DCO by adding a comment below using this text:


I have read the DCO document and I hereby sign the DCO.


You can retrigger this bot by commenting recheck in this Pull Request. Posted by the DCO Assistant Lite bot.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: install.sh waits 30s when the Linux gateway service fails to start

1 participant