Skip to content

Fix review findings of the cluster phase 4 integration (#137) - #163

Merged
viktor-ferenczi merged 28 commits into
phase4/quasar-integrationfrom
phase4/quasar-integration-fixes
Sep 21, 2026
Merged

viktor-ferenczi merged 28 commits into
phase4/quasar-integrationfrom
phase4/quasar-integration-fixes

Conversation

@viktor-ferenczi

@viktor-ferenczi viktor-ferenczi commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Fixes for the review of #137 at 3c07ea8. One commit per finding; the local ticket number (SE1-00xx) is in each commit body. Multi-Host behaviour is out of scope.

Fixed

Guided setup could never produce a working cluster

  • Executor tokens are named by Host ID; the host- prefix made every poll fail with executor_host_mismatch. A roster already written with the prefix is rewritten in place (SE1-0016).
  • The cluster spec emits administrators as SteamID64 numbers and rejects MaxPlayers < 2 before setup binds; the Gateway aborted at startup on strings (SE1-0017).
  • Local plugins without provenance metadata (UI-plugin companion DLLs) are left out of the cluster profile and named, instead of crashing -prepareManaged (SE1-0043, Quasar side; Magnetar side is fix(prepare): name the local plugin that lacks provenance metadata magnetar#59).
  • The failed-setup error is shown once (SE1-0044).

Quasar.Host stays up and recovers

  • Survives Gateway HTTP timeouts, per-cluster reconcile faults and a faulted Quasar tunnel (SE1-0022).
  • A reused PID is no longer a permanent unmanaged conflict: launch records carry boot ID + start ticks, final-state records are not inspected (SE1-0021).
  • A spawn cancelled by the lease budget no longer leaves a permanent Launching record; a started process whose identity cannot be committed is stopped and recorded as Failed (SE1-0023).
  • Invalid attachment / Gateway spec files are ignored; an unrecoverable activation or restore pauses only its cluster instead of exit code 2 on every start (SE1-0037).
  • Failed Gateway respawns back off 5 s … 5 min (SE1-0031, partial).

Quasar cluster management cannot be wedged by one bad file or one busy cluster

  • An unreadable operation record is quarantined as .corrupt and logged; the store recovers after a transient write failure; atomic writes are fsynced; store failures answer with the JSON envelope (SE1-0020).
  • Local operations are closed on every failure path and failed as interrupted_by_restart at startup, so they no longer block cluster deletion; unhandled InvalidOperationException on cluster routes is a JSON 409 (SE1-0026).
  • The reconciler, pending-operation and update loops skip a cluster whose lifecycle gate is busy (SE1-0027).
  • Unacknowledged Gateway requests expire after 10 minutes and can be withdrawn with DELETE …/operations/{operationId} (SE1-0035).
  • A torn cluster-credentials.json or host.json no longer aborts startup or breaks tunnel authentication for other Hosts (SE1-0037).
  • Stuck update, partial (SE1-0019): failed lifecycle shutdown is retried under a new key, Stopping/Starting record LastError when overdue, and POST …/update/abandon plus an Abandon update button unlock the lifecycle tools (refused while Hosts are half-activated).
  • Scheduled cluster backups, partial (SE1-0034): failures are logged and retried, one cluster or one corrupt backup.json no longer breaks the rest, and the panel says they need a stopped cluster.

Smaller

  • server.json goalState is numeric again so a downgrade still sees standalone servers; the cluster API keeps names (SE1-0036).
  • QUASAR_HOST_BINARY lets Host enrollment work from a development build; documented (SE1-0041).
  • Predefined DS worlds are listed once when a Steam and a managed DS install both exist (SE1-0042).
  • Removed committed .playwright-mcp/ artifacts (SE1-0038).

Found by the live re-test

  • Guided setup did not treat InvalidDataException as a failure: the status stayed on its last phase and the exception ended the operator's Blazor circuit (SE1-0045).
  • The web worker ignored SIGTERM while a Host was connected: the Host tunnel WebSockets held Kestrel's graceful shutdown until the 30 minute ShutdownTimeout. They now end on ApplicationStopping (SE1-0039).

Found by running a managed cluster with a headless client

  • Guided setup sent Magnetar's game version as binaryVersion; the Registry compares the node's Sandbox.Game.dll assembly version, so every node was rejected and the cluster never left Bootstrapping (SE1-0049).
  • A managed Gateway is Steam-only. QUASAR_CLUSTER_TEST_DIRECT_LISTEN / QUASAR_CLUSTER_TEST_DISABLE_STEAM (test only, private addresses only) add the Direct Transport frontend and drop Steam in the generated specification; needs CometWorks/cluster#18 (SE1-0048).
  • The Steam frontend needs Valve's steamclient.so under the Host account and nothing provided it: guided setup and conversion now send the copy from Quasar's managed SteamCMD to the Gateway Host, which stages it with a checksum and an atomic rename (SE1-0048).

Testing without GitHub releases (SE1-0046)

  • QUASAR_CLUSTER_ARCHIVE_URL / Quasar:ManagedRuntime:ClusterArchiveUrl points a test instance at a locally built ClusterForLinux-<version>.tar.gz (file:// or HTTP(S)) with its SHA256SUMS next to it. Verification is unchanged, the GitHub token is never sent there, and such a package is accepted only while the override is set.
  • QUASAR_MAGNETAR_ARCHIVE_URL also accepts file://; a rebuilt local archive is installed again. Both are documented in Docs/BuildingAndDevelopment.md.

Testing

  • dotnet build Quasar.sln -c Release: 0 errors, no new warnings in changed files.
  • Quasar.Tests: 411 passed (385 before; 26 new tests, one per fix where practical).
  • Quasar.Host --self-test: ok, extended for token names and roster migration, process identity, invalid state files, respawn backoff.
  • Live on the review install (web worker on 127.0.0.1:18080, Host local1 started by hand):
    • managed-credentials route writes token name local1;
    • zero-length file in Operations/Clusters → /api/ready 200, file renamed .corrupt, path logged;
    • Host attached to a listener that never answers survives repeated 15 s timeouts;
    • Gone launch record whose PID is held by an unrelated process → recovery-readiness 200 (was 409), same for a Running record from another boot;
    • /world-templates shows one row per predefined world.
  • Re-test after cluster v1.1.3 was released: guided setup ran to "Ready to start" on the review install. admission.json has numeric administrators, tokens.json names the Host ID, a failed setup shows one alert, a Gateway that exits at start is respawned with growing delays.
  • End to end with a headless client (locally built cluster releases from CometWorks/cluster#18, fed through QUASAR_CLUSTER_ARCHIVE_URL=file://…, Direct Transport only, no Steam on the Host account): the managed cluster reached Serving with two regular nodes and the World Authority, the Host's executor heartbeat was accepted, a headless Pulsar client from the cluster bench connected and spawned a living character (Gateway session Ready), server-measured movement was 23.6 m walking and 28.6 m / 42.4 m jetpack flight (up and strafe end at the ceiling and walls of the corridor the character spawns in), a rejoin re-possessed the same character, and goal Off shut the serving cluster down cleanly in 8 s. The steamclient.so Host route was exercised directly; a cluster serving through Steam was not started, and the Abandon update button was not seen in the browser.

Deliberately left for later

  • SE1-0018: resolved for now by cluster v1.1.3; the exact-bytes PluginSdk pin will block setup again at the next Magnetar release.
  • SE1-0019 remainder: timeouts that act, rollback, abandon during a half-applied activation, CLI command.
  • SE1-0024 clean-Down proof cleared by activation: needs a "never started since activation" proof in the writer-safety checks; analysed in the ticket, not attempted.
  • SE1-0031 remainder: verified-bundle hash cache, narrower Host execution gate. SE1-0034 remainder: scheduled backups for a 24/7 cluster.
  • SE1-0026 remainder: long operations still run on the HTTP request token. SE1-0035 remainder: no UI button for cancel.
  • SE1-0040 install-root Host/Port ignored: not reproducible at this head; no change.
  • New, open: goal Off does not stop the Host respawning a Gateway that never started; a fix is proposed in the ticket but contradicts an existing safety test, so it is not in this PR (SE1-0051).
  • New, open: guided setup accepts a world without a static grid and the converter then fails after the cluster ID is bound (SE1-0047).
  • Not started: SE1-0025, SE1-0028, SE1-0029, SE1-0030, SE1-0032 (identity half done with SE1-0021), SE1-0033.

The Gateway requires the polling Host ID to equal the token name, so the
host- prefix made every executor poll fail with executor_host_mismatch.
A roster already written with the prefix is rewritten in place because
it never authenticated.

SE1-0016
The Gateway reads admission.json strictly, so string administrators and
a member limit below 2 aborted it at startup. Administrators are now
validated as SteamID64 numbers and both values are checked before setup
binds or a conversion starts.

SE1-0017
An unreadable operation record is moved to .corrupt and named in the
log instead of disabling cluster management for every cluster. The
store recovers once its directory is writable again, atomic writes are
flushed to disk before the rename, and cluster routes without their own
handling answer store failures with the JSON error envelope.

SE1-0020
The HttpClient timeout raises TaskCanceledException, which was not in
the poll catch list, so an unreachable Gateway crashed the Host on every
start. Poll and Gateway reconcile failures are now contained per cluster
and a faulted Quasar tunnel is restarted instead of ending the executor
loop.

SE1-0022
Launch records now carry the kernel boot ID and the start time in ticks
since boot. A live PID with another identity proves the recorded process
is gone, so it reads as Missing instead of a permanent conflict, also
across reboots and wall-clock steps. Final-state records (Failed, Gone,
Stopped) are written only after the process was verified gone and are
no longer inspected. Records without an identity keep the conservative
start-time check on Linux.

SE1-0021
Cancellation is checked before the Launching record is written, not
between the record and the process start, so an executor lease that
runs out during bundle verification no longer leaves a permanent
'launch identity was not committed' conflict. A started process whose
identity cannot be committed is stopped through its handle and recorded
as Failed; the conflict remains only when it cannot be stopped.

SE1-0023
Host: an invalid attachment or Gateway spec file is ignored with a
message, and an activation or restore that cannot be recovered pauses
only its own cluster until the operation is retried, instead of ending
the Host with exit code 2 on every start.
Quasar: a torn cluster credential file is set aside rather than thrown
from a singleton constructor, and one unreadable host.json no longer
breaks tunnel authentication for every other Host.

SE1-0037
A local operation was finalized only for four exception types; any
other exception, a client disconnect or a Quasar restart left it
Running and blocked cluster deletion for good. Every failure now closes
the record, local operations still Running at startup are failed as
interrupted_by_restart, and cluster routes answer an unhandled
InvalidOperationException (a goal change during an update) with a JSON
409 instead of an HTML 500.

SE1-0026
The enum-level string converter made server.json store "On"/"Off",
which released Quasar versions cannot read, hiding standalone servers
after a downgrade. The enum now writes its number and reads both shapes;
the cluster API and cluster.json opt into names per property.

SE1-0036
The reconciler, the pending Gateway operation loop and the update loop
awaited each cluster's lifecycle gate in turn, so one backup, restore or
activation (held for up to two hours) stalled every other cluster. The
loops now take the gate without waiting and revisit a busy cluster on a
later pass.

SE1-0027
A Gateway that kept failing was respawned on every poll, and each
attempt verifies the whole bundle while the Host execution gate is held.
Respawns of the same spec are now spaced 5 s, 10 s, 20 s ... up to five
minutes; a changed spec or a new start generation retries at once and a
Gateway that stayed up for two minutes resets the delay.

SE1-0031
A shutdown, Gateway restart or kick submitted during an outage was
retried every two seconds forever and delivered whenever connectivity
returned. A request the Gateway never acknowledged now fails with
gateway_delivery_expired after ten minutes and can be withdrawn through
DELETE /api/v1/clusters/{id}/operations/{operationId}; an accepted one
cannot. An unchanged retry outcome is no longer rewritten to disk on
every pass.

SE1-0035
Partial fix; the full redesign (rollback, timeouts that act) stays open.
- The reconciler's lifecycle shutdown was keyed by lifecycle ID and
  terminal once Failed, so a failed drain was replayed every two seconds
  and the update waited for the clean-Down proof forever. A failed
  attempt is now retried under a new key after 60 s.
- Stopping (20 min) and Starting (30 min) record LastError once they are
  overdue, so the UI and API say what the update waits for.
- POST /api/v1/clusters/{id}/update/abandon and an 'Abandon update'
  button close a stuck update and unlock the lifecycle tools without
  rolling back; refused while Hosts are half-activated.

SE1-0019
Partial fix. Scheduled cluster backups are still offline backups; the
panel now says so instead of leaving a 24/7 cluster silently without
them. A failed scheduled capture is logged and retried after an hour
under a new operation key (same snapshot ID), one failing cluster no
longer aborts the loop for the others, and one unreadable backup.json
no longer breaks listing, scheduling and retention.

SE1-0034
Plain build output has no Host/Quasar.Host, so enrollment under
'dotnet run' could not work. QUASAR_HOST_BINARY points the web worker at
a published single-file Host; the error names it and the build docs show
the publish command. Releases keep using the packaged Host.

SE1-0041
A configured Steam DS install and the managed install ship the same
worlds, so every predefined world appeared twice. Templates are now
de-duplicated by category and relative path; the first install root wins.

SE1-0042
Guided setup deploys Quasar UI-plugin companion DLLs into the
preparation's Local folder, but writes provenance metadata only for its
own Agent. Magnetar -prepareManaged then crashed with a raw
FileNotFoundException on the first companion. Local plugins without
<name>.xml / <name>.dll.xml are now removed from the cluster profile
before preparation and named in the log and the setup phase.

SE1-0043 (Quasar side)
…ackup retry

SE1-0019 SE1-0026 SE1-0034 SE1-0035
Guided setup's failure filter did not include InvalidDataException,
which the world converter, the Magnetar export check and the Host
identity checks throw. The setup status stayed on its last phase with
no error and the exception ended the operator's Blazor circuit. Found
live: the converter refused the Empty World template (no static grid).

SE1-0045
A connected Host keeps its control WebSocket open for as long as it
runs. RequestAborted does not fire for a WebSocket at graceful shutdown,
so Kestrel waited for it until the 30 minute ShutdownTimeout: the web
worker logged 'Application is shutting down' on SIGTERM and kept
running. Found live: the hung worker exited within a second of the Host
disconnecting. The control and tunnel sockets now also end on
ApplicationStopping, like the Agent socket already does.

SE1-0039
QUASAR_CLUSTER_ARCHIVE_URL (Quasar:ManagedRuntime:ClusterArchiveUrl)
points a test instance at a ClusterForLinux-<version>.tar.gz served over
HTTP(S) with its SHA256SUMS next to it, mirroring the existing
QUASAR_MAGNETAR_ARCHIVE_URL. Verification is unchanged, the GitHub token
is never sent to that URL, and a package staged this way (release ID 0)
is accepted only while the override is set. Both overrides are now
documented in BuildingAndDevelopment.md.

SE1-0046
QUASAR_CLUSTER_ARCHIVE_URL and QUASAR_MAGNETAR_ARCHIVE_URL can point
straight at a locally built archive, so testing needs no temporary web
server. The cluster archive is verified as before against the
SHA256SUMS in the same directory. A local Magnetar archive is
identified by its size and modification time, so a rebuilt file is
installed again without changing the URL.

SE1-0046
…tend

The Gateway's Steam frontend loads Valve's steamclient.so from
~/.steam/sdk64. The cluster release cannot ship it and nothing provided
it, so on a Host account without Steam or SteamCMD the Gateway exited at
start with 'Steam GameServer.Init failed'. Guided setup and conversion
now send the copy from Quasar's managed SteamCMD to the Gateway Host,
which stages it with a checksum and an atomic rename.

SE1-0048
A managed Gateway is Steam-only, so no headless client could join a
Quasar-managed cluster. QUASAR_CLUSTER_TEST_DIRECT_LISTEN adds
directListen to the generated specification (loopback or private
addresses only: that frontend admits every client), and
QUASAR_CLUSTER_TEST_DISABLE_STEAM drops steamListen so the Gateway Host
needs no Steam client library, like the cluster bench. Needs a cluster
release that knows directListen.

SE1-0048
Guided setup put Magnetar's game version (1210014) into the
specification's binaryVersion. The Registry compares it with the node's
MySandboxGame.BuildVersion, the assembly version of Sandbox.Game.dll
(0.1.1.0), so every node of a guided cluster was rejected with
'Node admission rejected: binary version' and the cluster never left
Bootstrapping. The value now comes from the frozen Dedicated Server.

SE1-0049
@viktor-ferenczi
viktor-ferenczi merged commit ef65520 into phase4/quasar-integration Sep 21, 2026
@viktor-ferenczi
viktor-ferenczi deleted the phase4/quasar-integration-fixes branch September 21, 2026 00:10
viktor-ferenczi added a commit to CometWorks/magnetar that referenced this pull request Sep 21, 2026
`-prepareManaged` looked for `<dll>.xml` / `<name>.xml` next to a
`LocalPlugin` and passed the path straight to `RefuseLink`, which calls
`File.GetAttributes`. For a local plugin without metadata this raised an
unhandled `FileNotFoundException` (reported from Quasar guided cluster
setup with a UI-plugin companion DLL in `Local/`).

The path is now checked first and the existing
`InvalidDataException("Local plugin needs GitHubPlugin provenance
metadata: <id>")` is thrown. The message names the plugin ID and both
metadata paths that were tried (`<name>.xml` and `<name>.dll.xml`). The
sibling error for a metadata file that exists but is not a
`GitHubPlugin` now names that file too.

Testing: `dotnet build Legacy/Legacy.csproj -c Release` succeeds; no
unit test covers `ManagedPreparation` yet, so this is compile-checked
only.

Quasar side (leaves such plugins out of the cluster profile before
preparation) is CometWorks/quasar#163. Local ticket: SE1-0043.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant