Skip to content

Integrate cluster phase 4 - #137

Open
OwendB1 wants to merge 49 commits into
mainfrom
phase4/quasar-integration
Open

OwendB1 wants to merge 49 commits into
mainfrom
phase4/quasar-integration

Conversation

@OwendB1

@OwendB1 OwendB1 commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Integrates the cluster release into Quasar as a managed deployment: operators can provision a cluster, control its lifecycle, manage shared plugin configuration, convert existing servers, and perform backups, restore and full-downtime updates through the UI, API and CLI.

RBAC hardening was merged separately in #138. This PR covers Phase 4 integration and applies the existing authorization policies to its new cluster operations.

Deployment and lifecycle

  • Consume the complete verified upstream release as shipped, including node/World Authority plugins, Gateway, world converter and CLI. Cluster role plugins are not resolved through MagnetarHub.
  • Freeze Dedicated Server, Magnetar, transport, common/compatibility plugins and native assets into verified dependency snapshots. Bind those inputs and canonical configuration to a deployment revision; verify inventories, hashes and capabilities before Host activation.
  • Prepare all Hosts before activation and execute the Registry's node plan through authenticated, host-scoped executor leases. Fence reports, process adoption and kills by physical incarnation; preserve durable operation outcomes across retries and lost responses.
  • Start Gateway first and stop it last. Require authoritative clean-Down proof for normal teardown, activation and updates. Provide explicit stopped-fleet recovery for dirty shutdowns without wiping world or Registry data.

Operator workflows

  • Add cluster creation/import, fleet and Host views, lifecycle controls, diagnostics, maintenance, handover settings and operation progress, with corresponding API and CLI commands.
  • Deploy the same common plugins and canonical SDK configuration to every regular, World Authority, spare and replacement node. Shared configuration edits prepare a candidate revision for full-downtime activation; individual node edits are rejected.
  • Correlate optional Agent telemetry with the active deployment and exact runtime incarnation. Registry retains authority over placement, admission and lifecycle; running clusters do not depend on Quasar or Agent availability.
  • Add Convert to cluster and Convert to standalone server workflows. Conversion creates a separate stopped destination, preserves source data and backups, and records immutable requests for progress tracking and resume. Forward conversion uses reviewed plugin/configuration snapshots; reverse conversion assembles verified stopped-Host copies without requiring shared storage.
  • Add native multi-Host snapshots, verified world export retrieval, scheduled backups while cleanly stopped, and stopped-fleet restore with fresh credentials and stale-writer fencing.
  • Add resumable full-downtime update/rollback: preflight every Host, stop cleanly, activate, restart and verify convergence. Retain previous installations and reject incompatible storage formats.

Contracts and upstream dependencies

The Gateway admin/executor contract is vendored with source provenance; this PR does not require the obsolete private contract package.

Managed readiness and shared plugin services require coordinated upstream changes in cluster #12 and Magnetar #57. Published cluster 1.0.3 and Magnetar 2.4.2.0 do not provide these capabilities. Deployment checks require compatible release capabilities and the exact packaged PluginSdk hash.

Shared-state/node-aware SDK services are an upstream implementation, with Quasar providing packaging, configuration and observations. They remain a required part of the first post-beta release.

Validation

For the pushed implementation at 8edd33a:

  • Quasar test suite: 333 passed.
  • Quasar and Quasar.Host builds passed.
  • Offline Host preparation/import/tamper checks passed.
  • Regression coverage includes lifecycle fencing and recovery, lost-response replay, partial activation, update/rollback, conversion safety, snapshot assembly, corrupt artifact rejection, Agent identity and cluster route authorization.

See the seven-stage integration plan, verification record, conversion guide and plugin integration guide.

Acceptance and limits

Full local packaged-cluster acceptance remains open: live conversion parity, player cutover, outage behavior, updates and the complete P4/upstream regression scenarios are not established by offline tests. No Quasar web service or live game cluster was launched for the checks above. Existing game-reference and NuGet advisory warnings remain.

Conversion blocks unsupported access restrictions and non-SDK configuration providers. Reverse conversion exports canonical plugin settings for review/reapplication; arbitrary private files and shared plugin stores remain in backups and require explicit migration. Storage-format changes require an explicit migration and cannot be silently rolled back.

Coordinated upstream releases and live acceptance remain release gates. The validation above describes the pushed PR; ongoing local review fixes are excluded until pushed and reverified.

OwendB1 added 9 commits August 8, 2026 23:13
Write mutations before dispatch so dark-factory callers can replay or poll safely after Quasar restarts.
Register executor capacity without fabricating process state. Actualization waits for verified bundles and durable launch records.
Use write-ahead launch identity and node-authored readiness receipts so executor restarts cannot fabricate readiness or kill unknown processes.
Keep control credentials in process environments while Quasar durably records idempotent attachment mutations. Restrict the Host listener to loopback and reject bundle symlinks before launch.
Keep standalone controls unchanged while Gateway and Host truth populate the shared server views, cluster details, and Hosts tool.
Persist level-triggered Gateway specs in Quasar.Host, re-adopt exact processes after restart, and expose the same durable headless API with a clean-shutdown Off gate.
Persist cluster goals and Gateway specs, route UI/API changes through durable operations, and sequence Gateway-first start and clean stop in normal and headless modes.
Expose versioned cluster reads and durable lifecycle mutations through the packaged launcher with environment-only credentials, stable exit codes, and resumable operation wait.
Viewer sessions could reach mutating Blazor controls and retain stale role claims after RBAC changes. Revalidate roles at request and action boundaries, restrict mutation routes, disable subnet admin bypass by default, and prevent last-admin removal.

Existing installs with AllowSameSubnet enabled must disable it explicitly.
@OwendB1

OwendB1 commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

Closed because this branch also contains Phase 4 integration work. RBAC-only replacement: #138.

@OwendB1 OwendB1 closed this Aug 11, 2026
@OwendB1 OwendB1 reopened this Aug 11, 2026
Retain cluster execution and current main fixes, with upstream clustering
plans taking priority over earlier integration-only decisions.

Alias the Bootstrap test reference to disambiguate shared networking types.
Record the corrected continuation plan and baseline validation.

Validation: solution build, 192 tests, executor self-test, and foreground
and service Bootstrap shutdown checks.
Bind telemetry and plugin edits to their originating incarnation and connection. Reject late messages after replacement so stale state cannot cross into a new node.
- Extend admin/executor contracts for handover and deployment readiness
- Document cluster plugin management, conversion, and integration verification
@OwendB1 OwendB1 changed the title Integrate Phase 4 and harden RBAC Integrate cluster phase 4 and harden RBAC Sep 20, 2026
@OwendB1 OwendB1 changed the title Integrate cluster phase 4 and harden RBAC Integrate cluster phase 4 Sep 20, 2026
- Capture and convert SDK-tracked configuration types safely
- Harden restore, drift, persistence, and unmanaged migration guidance
- Expand regression coverage and integration verification
- Include Linux and Windows Host archives in releases and checksums
- Document Host deployment and managed cluster behavior
- Prevent Python bytecode writes and correct dependency validation
- Retain console events during Gateway outages with scope enforcement
- Complete shutdown journals from verified clean-stop proof
- Rework cluster cards and deployment sections for responsive UI
- Keep deployment configuration on the cluster details page
- Update related documentation
- Keep configuration destinations available through the navigation bar
- Update workflow documentation to match
@OwendB1
OwendB1 marked this pull request as ready for review September 20, 2026 11:12
- Add machine enrollment, credential provisioning, and Host control tunnels
- Document guided setup, deletion, prerequisites, and acceptance checks
- Embed native Host libraries in single-file publishing
- Preserve actionable package errors through the release API
- Add regression coverage for HTTP, network, timeout, and cleanup failures
Fixes for the review of #137 at `3c07ea8`. One commit per finding; the
local ticket number (SE1-00xx) is in each commit body. Multi-Host
behaviour is out of scope.

## Fixed

**Guided setup could never produce a working cluster**
- Executor tokens are named by Host ID; the `host-` prefix made every
poll fail with `executor_host_mismatch`. A roster already written with
the prefix is rewritten in place (SE1-0016).
- The cluster spec emits administrators as SteamID64 numbers and rejects
`MaxPlayers < 2` before setup binds; the Gateway aborted at startup on
strings (SE1-0017).
- Local plugins without provenance metadata (UI-plugin companion DLLs)
are left out of the cluster profile and named, instead of crashing
`-prepareManaged` (SE1-0043, Quasar side; Magnetar side is
CometWorks/magnetar#59).
- The failed-setup error is shown once (SE1-0044).

**Quasar.Host stays up and recovers**
- Survives Gateway HTTP timeouts, per-cluster reconcile faults and a
faulted Quasar tunnel (SE1-0022).
- A reused PID is no longer a permanent unmanaged conflict: launch
records carry boot ID + start ticks, final-state records are not
inspected (SE1-0021).
- A spawn cancelled by the lease budget no longer leaves a permanent
`Launching` record; a started process whose identity cannot be committed
is stopped and recorded as Failed (SE1-0023).
- Invalid attachment / Gateway spec files are ignored; an unrecoverable
activation or restore pauses only its cluster instead of exit code 2 on
every start (SE1-0037).
- Failed Gateway respawns back off 5 s … 5 min (SE1-0031, partial).

**Quasar cluster management cannot be wedged by one bad file or one busy
cluster**
- An unreadable operation record is quarantined as `.corrupt` and
logged; the store recovers after a transient write failure; atomic
writes are fsynced; store failures answer with the JSON envelope
(SE1-0020).
- Local operations are closed on every failure path and failed as
`interrupted_by_restart` at startup, so they no longer block cluster
deletion; unhandled `InvalidOperationException` on cluster routes is a
JSON 409 (SE1-0026).
- The reconciler, pending-operation and update loops skip a cluster
whose lifecycle gate is busy (SE1-0027).
- Unacknowledged Gateway requests expire after 10 minutes and can be
withdrawn with `DELETE …/operations/{operationId}` (SE1-0035).
- A torn `cluster-credentials.json` or `host.json` no longer aborts
startup or breaks tunnel authentication for other Hosts (SE1-0037).
- Stuck update, partial (SE1-0019): failed lifecycle shutdown is retried
under a new key, Stopping/Starting record `LastError` when overdue, and
`POST …/update/abandon` plus an **Abandon update** button unlock the
lifecycle tools (refused while Hosts are half-activated).
- Scheduled cluster backups, partial (SE1-0034): failures are logged and
retried, one cluster or one corrupt `backup.json` no longer breaks the
rest, and the panel says they need a stopped cluster.

**Smaller**
- `server.json` `goalState` is numeric again so a downgrade still sees
standalone servers; the cluster API keeps names (SE1-0036).
- `QUASAR_HOST_BINARY` lets Host enrollment work from a development
build; documented (SE1-0041).
- Predefined DS worlds are listed once when a Steam and a managed DS
install both exist (SE1-0042).
- Removed committed `.playwright-mcp/` artifacts (SE1-0038).

**Found by the live re-test**
- Guided setup did not treat `InvalidDataException` as a failure: the
status stayed on its last phase and the exception ended the operator's
Blazor circuit (SE1-0045).
- The web worker ignored SIGTERM while a Host was connected: the Host
tunnel WebSockets held Kestrel's graceful shutdown until the 30 minute
`ShutdownTimeout`. They now end on `ApplicationStopping` (SE1-0039).

**Found by running a managed cluster with a headless client**
- Guided setup sent Magnetar's game version as `binaryVersion`; the
Registry compares the node's `Sandbox.Game.dll` assembly version, so
every node was rejected and the cluster never left Bootstrapping
(SE1-0049).
- A managed Gateway is Steam-only. `QUASAR_CLUSTER_TEST_DIRECT_LISTEN` /
`QUASAR_CLUSTER_TEST_DISABLE_STEAM` (test only, private addresses only)
add the Direct Transport frontend and drop Steam in the generated
specification; needs CometWorks/cluster#18 (SE1-0048).
- The Steam frontend needs Valve's `steamclient.so` under the Host
account and nothing provided it: guided setup and conversion now send
the copy from Quasar's managed SteamCMD to the Gateway Host, which
stages it with a checksum and an atomic rename (SE1-0048).

**Testing without GitHub releases** (SE1-0046)
- `QUASAR_CLUSTER_ARCHIVE_URL` /
`Quasar:ManagedRuntime:ClusterArchiveUrl` points a test instance at a
locally built `ClusterForLinux-<version>.tar.gz` (`file://` or HTTP(S))
with its `SHA256SUMS` next to it. Verification is unchanged, the GitHub
token is never sent there, and such a package is accepted only while the
override is set.
- `QUASAR_MAGNETAR_ARCHIVE_URL` also accepts `file://`; a rebuilt local
archive is installed again. Both are documented in
`Docs/BuildingAndDevelopment.md`.

## Testing
- `dotnet build Quasar.sln -c Release`: 0 errors, no new warnings in
changed files.
- `Quasar.Tests`: 411 passed (385 before; 26 new tests, one per fix
where practical).
- `Quasar.Host --self-test`: ok, extended for token names and roster
migration, process identity, invalid state files, respawn backoff.
- Live on the review install (web worker on `127.0.0.1:18080`, Host
`local1` started by hand):
  - managed-credentials route writes token name `local1`;
- zero-length file in `Operations/Clusters` → `/api/ready` 200, file
renamed `.corrupt`, path logged;
- Host attached to a listener that never answers survives repeated 15 s
timeouts;
- `Gone` launch record whose PID is held by an unrelated process →
`recovery-readiness` 200 (was 409), same for a `Running` record from
another boot;
  - `/world-templates` shows one row per predefined world.
- Re-test after cluster v1.1.3 was released: guided setup ran to "Ready
to start" on the review install. `admission.json` has numeric
administrators, `tokens.json` names the Host ID, a failed setup shows
one alert, a Gateway that exits at start is respawned with growing
delays.
- End to end with a headless client (locally built cluster releases from
CometWorks/cluster#18, fed through
`QUASAR_CLUSTER_ARCHIVE_URL=file://…`, Direct Transport only, no Steam
on the Host account): the managed cluster reached Serving with two
regular nodes and the World Authority, the Host's executor heartbeat was
accepted, a headless Pulsar client from the cluster bench connected and
spawned a living character (Gateway session `Ready`), server-measured
movement was 23.6 m walking and 28.6 m / 42.4 m jetpack flight (up and
strafe end at the ceiling and walls of the corridor the character spawns
in), a rejoin re-possessed the same character, and goal Off shut the
serving cluster down cleanly in 8 s. The `steamclient.so` Host route was
exercised directly; a cluster serving through Steam was not started, and
the Abandon update button was not seen in the browser.

## Deliberately left for later
- SE1-0018: resolved for now by cluster v1.1.3; the exact-bytes
PluginSdk pin will block setup again at the next Magnetar release.
- SE1-0019 remainder: timeouts that act, rollback, abandon during a
half-applied activation, CLI command.
- SE1-0024 clean-Down proof cleared by activation: needs a "never
started since activation" proof in the writer-safety checks; analysed in
the ticket, not attempted.
- SE1-0031 remainder: verified-bundle hash cache, narrower Host
execution gate. SE1-0034 remainder: scheduled backups for a 24/7
cluster.
- SE1-0026 remainder: long operations still run on the HTTP request
token. SE1-0035 remainder: no UI button for cancel.
- SE1-0040 install-root Host/Port ignored: not reproducible at this
head; no change.
- New, open: goal Off does not stop the Host respawning a Gateway that
never started; a fix is proposed in the ticket but contradicts an
existing safety test, so it is not in this PR (SE1-0051).
- New, open: guided setup accepts a world without a static grid and the
converter then fails after the cluster ID is bound (SE1-0047).
- Not started: SE1-0025, SE1-0028, SE1-0029, SE1-0030, SE1-0032
(identity half done with SE1-0021), SE1-0033.
- Notify operators of newer stable cluster packages
- Prevent developer checkouts from shadowing managed transport
…phase4/quasar-integration

# Conflicts:
#	Quasar.Tests/ClusterSetupTests.cs
#	Quasar/Services/ClusterSetupService.cs
- Re-select compatible releases and rebuild stale plugin exports
- Validate PluginSdk pins and malformed release metadata
- Preserve Host state during binary updates with rollback on restart failure
- Recover stale supervised workers without killing unrelated processes
- Install authenticated 15-minute update timers with rollback
- Support verified binary downloads and legacy local migration
- Require explicit preparation after importing edited JSON
- Show profile plugin names in deployment configuration
- Add release staging, dependency freezing, host preparation, and full-downtime update actions
- Persist prepared deployment candidates across catalog reloads
- Enforce clean-shutdown recovery and document the new workflow
- Prevent starts until Gateway stop state is verified
- Stop retrying failed shutdown after all nodes empty
- Clarify drain, push setup, and recovery status in UI and docs
- The staged diff is whitespace-clean and the migration paths appear consistent.
- I could not edit or run tests because the workspace is mounted read-only.
- Allow operators to start a fresh setup after failure and clear stale IDs when the registration is gone.
- Add the proposal and link it from the README.
- Link the README to the plan in the clustering-plan repository.
- Reviewed implementation and docs; `git diff --cached --check` passed.
- Focused tests could not run: the workspace is read-only, and MSBuild could not create its temp/cache directory.
- List update notices and manage browser push settings on Notifications.
- Link the bell directly to the current notice.
- Route cluster chat through fenced PluginSdk broadcasts and show cluster targets, monitoring, and controls in Quasar.
- Update setup and configuration documentation.
- Tests could not run: the read-only filesystem prevented MSBuild from creating temporary files.
- Add server selection, status, activity, and live preview controls.
- Include cluster presence without requiring channel bindings; document defaults and behavior.
- Follow the Quasar registration through activation and report worker failures clearly.
- Move Security, UI Plugins, and Backups into Tools.
- Update the navigation documentation to match
- Auto-open the Mods and plugins panel when changes are available; preserve manual toggles.
- Use a regular link for cluster details so it can open in a new tab.
- Add a saved-preparation editor for cluster-wide plugin settings
- Clarify cluster overview and maintenance navigation
- Update UI labels and related documentation references
- Point Brave users to the push messaging setting when registration fails.
- Add persistent queued mod and plugin update requests, with activation blocked until pinned content preparation is available.
- Simplify maintenance into staging and applying pending updates, including plugin configuration changes.
- Update configuration docs and focused tests.
- Save cluster display name and shutdown grace period.
- Stage compatible profiles and show planned node slots alongside registered nodes.
- Explain managed mod staging and world migration gaps
- Require verified full downtime before restaging active content
Hide expected Gateway connectivity errors while goal state is Off and bind prepared profile selections to the saved profile content. Managed node and World Authority launches force local mod loading so recycled nodes cannot fetch a newer Workshop revision. Modded deployments need staged local payloads before Host rollout.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants