Conversation
Write mutations before dispatch so dark-factory callers can replay or poll safely after Quasar restarts.
Register executor capacity without fabricating process state. Actualization waits for verified bundles and durable launch records.
Use write-ahead launch identity and node-authored readiness receipts so executor restarts cannot fabricate readiness or kill unknown processes.
Keep control credentials in process environments while Quasar durably records idempotent attachment mutations. Restrict the Host listener to loopback and reject bundle symlinks before launch.
Keep standalone controls unchanged while Gateway and Host truth populate the shared server views, cluster details, and Hosts tool.
Persist level-triggered Gateway specs in Quasar.Host, re-adopt exact processes after restart, and expose the same durable headless API with a clean-shutdown Off gate.
Persist cluster goals and Gateway specs, route UI/API changes through durable operations, and sequence Gateway-first start and clean stop in normal and headless modes.
Expose versioned cluster reads and durable lifecycle mutations through the packaged launcher with environment-only credentials, stable exit codes, and resumable operation wait.
Viewer sessions could reach mutating Blazor controls and retain stale role claims after RBAC changes. Revalidate roles at request and action boundaries, restrict mutation routes, disable subnet admin bypass by default, and prevent last-admin removal. Existing installs with AllowSameSubnet enabled must disable it explicitly.
Contributor
Author
|
Closed because this branch also contains Phase 4 integration work. RBAC-only replacement: #138. |
Retain cluster execution and current main fixes, with upstream clustering plans taking priority over earlier integration-only decisions. Alias the Bootstrap test reference to disambiguate shared networking types. Record the corrected continuation plan and baseline validation. Validation: solution build, 192 tests, executor self-test, and foreground and service Bootstrap shutdown checks.
Bind telemetry and plugin edits to their originating incarnation and connection. Reject late messages after replacement so stale state cannot cross into a new node.
- Extend admin/executor contracts for handover and deployment readiness - Document cluster plugin management, conversion, and integration verification
- Capture and convert SDK-tracked configuration types safely - Harden restore, drift, persistence, and unmanaged migration guidance - Expand regression coverage and integration verification
- Include Linux and Windows Host archives in releases and checksums - Document Host deployment and managed cluster behavior - Prevent Python bytecode writes and correct dependency validation
- Retain console events during Gateway outages with scope enforcement - Complete shutdown journals from verified clean-stop proof - Rework cluster cards and deployment sections for responsive UI
- Keep deployment configuration on the cluster details page - Update related documentation
- Keep configuration destinations available through the navigation bar - Update workflow documentation to match
OwendB1
marked this pull request as ready for review
September 20, 2026 11:12
- Add machine enrollment, credential provisioning, and Host control tunnels - Document guided setup, deletion, prerequisites, and acceptance checks - Embed native Host libraries in single-file publishing
- Preserve actionable package errors through the release API - Add regression coverage for HTTP, network, timeout, and cleanup failures
Fixes for the review of #137 at `3c07ea8`. One commit per finding; the local ticket number (SE1-00xx) is in each commit body. Multi-Host behaviour is out of scope. ## Fixed **Guided setup could never produce a working cluster** - Executor tokens are named by Host ID; the `host-` prefix made every poll fail with `executor_host_mismatch`. A roster already written with the prefix is rewritten in place (SE1-0016). - The cluster spec emits administrators as SteamID64 numbers and rejects `MaxPlayers < 2` before setup binds; the Gateway aborted at startup on strings (SE1-0017). - Local plugins without provenance metadata (UI-plugin companion DLLs) are left out of the cluster profile and named, instead of crashing `-prepareManaged` (SE1-0043, Quasar side; Magnetar side is CometWorks/magnetar#59). - The failed-setup error is shown once (SE1-0044). **Quasar.Host stays up and recovers** - Survives Gateway HTTP timeouts, per-cluster reconcile faults and a faulted Quasar tunnel (SE1-0022). - A reused PID is no longer a permanent unmanaged conflict: launch records carry boot ID + start ticks, final-state records are not inspected (SE1-0021). - A spawn cancelled by the lease budget no longer leaves a permanent `Launching` record; a started process whose identity cannot be committed is stopped and recorded as Failed (SE1-0023). - Invalid attachment / Gateway spec files are ignored; an unrecoverable activation or restore pauses only its cluster instead of exit code 2 on every start (SE1-0037). - Failed Gateway respawns back off 5 s … 5 min (SE1-0031, partial). **Quasar cluster management cannot be wedged by one bad file or one busy cluster** - An unreadable operation record is quarantined as `.corrupt` and logged; the store recovers after a transient write failure; atomic writes are fsynced; store failures answer with the JSON envelope (SE1-0020). - Local operations are closed on every failure path and failed as `interrupted_by_restart` at startup, so they no longer block cluster deletion; unhandled `InvalidOperationException` on cluster routes is a JSON 409 (SE1-0026). - The reconciler, pending-operation and update loops skip a cluster whose lifecycle gate is busy (SE1-0027). - Unacknowledged Gateway requests expire after 10 minutes and can be withdrawn with `DELETE …/operations/{operationId}` (SE1-0035). - A torn `cluster-credentials.json` or `host.json` no longer aborts startup or breaks tunnel authentication for other Hosts (SE1-0037). - Stuck update, partial (SE1-0019): failed lifecycle shutdown is retried under a new key, Stopping/Starting record `LastError` when overdue, and `POST …/update/abandon` plus an **Abandon update** button unlock the lifecycle tools (refused while Hosts are half-activated). - Scheduled cluster backups, partial (SE1-0034): failures are logged and retried, one cluster or one corrupt `backup.json` no longer breaks the rest, and the panel says they need a stopped cluster. **Smaller** - `server.json` `goalState` is numeric again so a downgrade still sees standalone servers; the cluster API keeps names (SE1-0036). - `QUASAR_HOST_BINARY` lets Host enrollment work from a development build; documented (SE1-0041). - Predefined DS worlds are listed once when a Steam and a managed DS install both exist (SE1-0042). - Removed committed `.playwright-mcp/` artifacts (SE1-0038). **Found by the live re-test** - Guided setup did not treat `InvalidDataException` as a failure: the status stayed on its last phase and the exception ended the operator's Blazor circuit (SE1-0045). - The web worker ignored SIGTERM while a Host was connected: the Host tunnel WebSockets held Kestrel's graceful shutdown until the 30 minute `ShutdownTimeout`. They now end on `ApplicationStopping` (SE1-0039). **Found by running a managed cluster with a headless client** - Guided setup sent Magnetar's game version as `binaryVersion`; the Registry compares the node's `Sandbox.Game.dll` assembly version, so every node was rejected and the cluster never left Bootstrapping (SE1-0049). - A managed Gateway is Steam-only. `QUASAR_CLUSTER_TEST_DIRECT_LISTEN` / `QUASAR_CLUSTER_TEST_DISABLE_STEAM` (test only, private addresses only) add the Direct Transport frontend and drop Steam in the generated specification; needs CometWorks/cluster#18 (SE1-0048). - The Steam frontend needs Valve's `steamclient.so` under the Host account and nothing provided it: guided setup and conversion now send the copy from Quasar's managed SteamCMD to the Gateway Host, which stages it with a checksum and an atomic rename (SE1-0048). **Testing without GitHub releases** (SE1-0046) - `QUASAR_CLUSTER_ARCHIVE_URL` / `Quasar:ManagedRuntime:ClusterArchiveUrl` points a test instance at a locally built `ClusterForLinux-<version>.tar.gz` (`file://` or HTTP(S)) with its `SHA256SUMS` next to it. Verification is unchanged, the GitHub token is never sent there, and such a package is accepted only while the override is set. - `QUASAR_MAGNETAR_ARCHIVE_URL` also accepts `file://`; a rebuilt local archive is installed again. Both are documented in `Docs/BuildingAndDevelopment.md`. ## Testing - `dotnet build Quasar.sln -c Release`: 0 errors, no new warnings in changed files. - `Quasar.Tests`: 411 passed (385 before; 26 new tests, one per fix where practical). - `Quasar.Host --self-test`: ok, extended for token names and roster migration, process identity, invalid state files, respawn backoff. - Live on the review install (web worker on `127.0.0.1:18080`, Host `local1` started by hand): - managed-credentials route writes token name `local1`; - zero-length file in `Operations/Clusters` → `/api/ready` 200, file renamed `.corrupt`, path logged; - Host attached to a listener that never answers survives repeated 15 s timeouts; - `Gone` launch record whose PID is held by an unrelated process → `recovery-readiness` 200 (was 409), same for a `Running` record from another boot; - `/world-templates` shows one row per predefined world. - Re-test after cluster v1.1.3 was released: guided setup ran to "Ready to start" on the review install. `admission.json` has numeric administrators, `tokens.json` names the Host ID, a failed setup shows one alert, a Gateway that exits at start is respawned with growing delays. - End to end with a headless client (locally built cluster releases from CometWorks/cluster#18, fed through `QUASAR_CLUSTER_ARCHIVE_URL=file://…`, Direct Transport only, no Steam on the Host account): the managed cluster reached Serving with two regular nodes and the World Authority, the Host's executor heartbeat was accepted, a headless Pulsar client from the cluster bench connected and spawned a living character (Gateway session `Ready`), server-measured movement was 23.6 m walking and 28.6 m / 42.4 m jetpack flight (up and strafe end at the ceiling and walls of the corridor the character spawns in), a rejoin re-possessed the same character, and goal Off shut the serving cluster down cleanly in 8 s. The `steamclient.so` Host route was exercised directly; a cluster serving through Steam was not started, and the Abandon update button was not seen in the browser. ## Deliberately left for later - SE1-0018: resolved for now by cluster v1.1.3; the exact-bytes PluginSdk pin will block setup again at the next Magnetar release. - SE1-0019 remainder: timeouts that act, rollback, abandon during a half-applied activation, CLI command. - SE1-0024 clean-Down proof cleared by activation: needs a "never started since activation" proof in the writer-safety checks; analysed in the ticket, not attempted. - SE1-0031 remainder: verified-bundle hash cache, narrower Host execution gate. SE1-0034 remainder: scheduled backups for a 24/7 cluster. - SE1-0026 remainder: long operations still run on the HTTP request token. SE1-0035 remainder: no UI button for cancel. - SE1-0040 install-root Host/Port ignored: not reproducible at this head; no change. - New, open: goal Off does not stop the Host respawning a Gateway that never started; a fix is proposed in the ticket but contradicts an existing safety test, so it is not in this PR (SE1-0051). - New, open: guided setup accepts a world without a static grid and the converter then fails after the cluster ID is bound (SE1-0047). - Not started: SE1-0025, SE1-0028, SE1-0029, SE1-0030, SE1-0032 (identity half done with SE1-0021), SE1-0033.
- Notify operators of newer stable cluster packages - Prevent developer checkouts from shadowing managed transport
…phase4/quasar-integration # Conflicts: # Quasar.Tests/ClusterSetupTests.cs # Quasar/Services/ClusterSetupService.cs
- Re-select compatible releases and rebuild stale plugin exports - Validate PluginSdk pins and malformed release metadata
- Preserve Host state during binary updates with rollback on restart failure - Recover stale supervised workers without killing unrelated processes
- Install authenticated 15-minute update timers with rollback - Support verified binary downloads and legacy local migration
- Require explicit preparation after importing edited JSON - Show profile plugin names in deployment configuration
- Add release staging, dependency freezing, host preparation, and full-downtime update actions - Persist prepared deployment candidates across catalog reloads - Enforce clean-shutdown recovery and document the new workflow
- Prevent starts until Gateway stop state is verified - Stop retrying failed shutdown after all nodes empty - Clarify drain, push setup, and recovery status in UI and docs
- The staged diff is whitespace-clean and the migration paths appear consistent. - I could not edit or run tests because the workspace is mounted read-only.
- Allow operators to start a fresh setup after failure and clear stale IDs when the registration is gone.
- Add the proposal and link it from the README.
- Link the README to the plan in the clustering-plan repository.
- Reviewed implementation and docs; `git diff --cached --check` passed. - Focused tests could not run: the workspace is read-only, and MSBuild could not create its temp/cache directory.
- List update notices and manage browser push settings on Notifications. - Link the bell directly to the current notice.
- Route cluster chat through fenced PluginSdk broadcasts and show cluster targets, monitoring, and controls in Quasar. - Update setup and configuration documentation. - Tests could not run: the read-only filesystem prevented MSBuild from creating temporary files.
- Add server selection, status, activity, and live preview controls. - Include cluster presence without requiring channel bindings; document defaults and behavior.
- Follow the Quasar registration through activation and report worker failures clearly. - Move Security, UI Plugins, and Backups into Tools.
- Update the navigation documentation to match
- Auto-open the Mods and plugins panel when changes are available; preserve manual toggles. - Use a regular link for cluster details so it can open in a new tab.
- Add a saved-preparation editor for cluster-wide plugin settings - Clarify cluster overview and maintenance navigation
- Update UI labels and related documentation references
- Point Brave users to the push messaging setting when registration fails.
- Add persistent queued mod and plugin update requests, with activation blocked until pinned content preparation is available. - Simplify maintenance into staging and applying pending updates, including plugin configuration changes. - Update configuration docs and focused tests.
- Save cluster display name and shutdown grace period. - Stage compatible profiles and show planned node slots alongside registered nodes.
- Explain managed mod staging and world migration gaps - Require verified full downtime before restaging active content
Hide expected Gateway connectivity errors while goal state is Off and bind prepared profile selections to the saved profile content. Managed node and World Authority launches force local mod loading so recycled nodes cannot fetch a newer Workshop revision. Modded deployments need staged local payloads before Host rollout.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Integrates the cluster release into Quasar as a managed deployment: operators can provision a cluster, control its lifecycle, manage shared plugin configuration, convert existing servers, and perform backups, restore and full-downtime updates through the UI, API and CLI.
RBAC hardening was merged separately in #138. This PR covers Phase 4 integration and applies the existing authorization policies to its new cluster operations.
Deployment and lifecycle
Operator workflows
Contracts and upstream dependencies
The Gateway admin/executor contract is vendored with source provenance; this PR does not require the obsolete private contract package.
Managed readiness and shared plugin services require coordinated upstream changes in cluster #12 and Magnetar #57. Published cluster 1.0.3 and Magnetar 2.4.2.0 do not provide these capabilities. Deployment checks require compatible release capabilities and the exact packaged PluginSdk hash.
Shared-state/node-aware SDK services are an upstream implementation, with Quasar providing packaging, configuration and observations. They remain a required part of the first post-beta release.
Validation
For the pushed implementation at
8edd33a:See the seven-stage integration plan, verification record, conversion guide and plugin integration guide.
Acceptance and limits
Full local packaged-cluster acceptance remains open: live conversion parity, player cutover, outage behavior, updates and the complete P4/upstream regression scenarios are not established by offline tests. No Quasar web service or live game cluster was launched for the checks above. Existing game-reference and NuGet advisory warnings remain.
Conversion blocks unsupported access restrictions and non-SDK configuration providers. Reverse conversion exports canonical plugin settings for review/reapplication; arbitrary private files and shared plugin stores remain in backups and require explicit migration. Storage-format changes require an explicit migration and cannot be silently rolled back.
Coordinated upstream releases and live acceptance remain release gates. The validation above describes the pushed PR; ongoing local review fixes are excluded until pushed and reverified.