operator: StorageSiteDeployment — deploy a managed site's storage from the hub - #620
Merged
schmidt-scaled merged 18 commits intoOct 3, 2026
Merged
Conversation
…t its end A volume's PV keeps the handle it was created with across fail-overs, and the chain of replication relationships behind that handle alternates between the sites: every fail-over adds a hop to the other side. resolveToLocalReplica walked the chain to its active end for every csi-addons Replication RPC. After an unplanned fail-over A->B, Ramen makes the old primary on A secondary: DemoteVolume (and DisableVolumeReplication on VR deletion) on site A resolved to the chain's end -- the NEW primary on site B -- and demoted it (live 2026-10-02, realbed WordPress: the demote fenced the live primary's paths and took demote snapshots of it; the VRG on A never became secondary, so the fail-back never got PeerReady). The driver needs to know which clusters are its own. The operator marks the clusters it manages `local` in the CSI secret's entries (an entry another site registered for cross-cluster handle resolution is not), and the resolver returns the chain's last member on a local cluster. A secret that marks no cluster local -- an operator predating the flag -- keeps the previous behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The old primary of an unplanned fail-over is reaped by the control plane once its fail-over completed; when the site returns Ramen still demotes it and deletes its VolumeReplication. Resolved to that member, the demote and the detach got a 404 and the VR stayed Degraded (live 2026-10-02, site A, 80e3e748). A 404 on a member of a chain the backend records is nothing left to demote or detach: success. A 404 on a handle without any relationship stays NotFound. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… the chain's active end After an unplanned fail-over the recovered site's old primary is reaped once its fail-over completed. Ramen still resyncs that site's secondary VR (and reads its status). Addressed to the reaped member both got a 404 and the VR stayed Degraded (live 2026-10-02, site A). sbcli's replication_failback is addressed to the failed-over clone and, given the original source cluster, re-aims the clone's replication at that site's node so only the delta ships: Resync falls back to the chain's active end with the local cluster as the fail-back source, and the status read follows. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ocal member The control plane's failover endpoint takes the volume that holds the data (the pairing's source) and creates the clone on the replication target. The local member on the promoting site is the volume being replaced -- demoted, or already reaped: on 2026-10-02 the fail-back promote on site A hit the reaped 80e3e748 and 404ed while the live primary on B held the data. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
sbcli's replication_failback takes the volume that holds the data (the failed-over clone) and, given the recovered site as source cluster, re-aims its replication at that site's node. Addressed to the local member -- the demoted or reaped old primary -- it configured nothing for the live clone: after the unplanned fail-over of Gitea the clones on A had no replication and lastGroupSyncTime stayed empty (2026-10-02). The pairing's status is the active end's for the same reason. Demote and Disable keep the local member; Promote already uses the active end. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…volume carries
A PersistentVolume without fsType -- a static PV adopting an existing
volume (the storage operator's test fail-over clone), or one Ramen restored
without the field -- gave the node plugin an empty fsType, which it turned
into "ext4" before the device was looked at; the filesystem layer then
refused every such XFS volume ("the volume carries xfs and this plan asks
for ext4"), and the bubble VM never started (realbed 2026-10-03).
An empty FsType now expresses no opinion: the layer mounts the filesystem
the device carries, formats a blank device as DefaultFsType (ext4), records
what it found as its params, and the node plugin remembers that type for
the volume. A named filesystem keeps refusing a device carrying another.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…umeMode A VM's disk is a Block claim. The drill's bubble PV and PVC omitted the mode, defaulted to Filesystem, and the kubelet asked the node plugin to mount the raw guest disk (mount: exit status 32); the bubble VM never started (realbed 2026-10-03). The source PV's volumeMode is recorded on the clone and set on both objects. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The console templates and their values block of feature/control-center-ui (PR #512), copied as they are: self-contained (sbcc.* helpers only), so the chart deploys the UI with the control plane on this line too. One addition: controlCenter.service.nodePort pins the node port of a NodePort Service. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… by 8 hops The PV keeps the original handle while every relocate and fail-over appends a clone, so the chain behind a volume grows by one per move and is never compacted. Both walks (the node plugin's redirect to the active volume and the controller's resolveChain) stopped after 8 hops; the ninth move of a volume -- an unplanned fail-over on 2026-10-03 -- left its clone one hop out of reach, the node fell back to the stashed context and attached the original on the partitioned site, and the VM never started. The walks now track the members visited and stop on a loop, with a bound of 256 as a guard far above any real chain. Tests: a 12-move chain alternating clusters, and a looping pair. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The hand-written sourceVolumeMode entry differed from the generated one in wrapping; the Manifests check keeps the three copies identical to the generator's output. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…erated TestFailover CRD Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
schmidt-scaled
pushed a commit
that referenced
this pull request
Oct 3, 2026
…loyed from the hub Disaster recovery -> Site storage lists the hub's StorageSiteDeployments (storage.simplyblock.io/v1alpha2, operator PR #620): one per managed site, with its phase, the nodes the discovery found and, once approved, the storage cluster with its id, pool and nodes. - "Deploy storage" picks a managed cluster (ManagedClusters, else the site profiles), the discovery scope and the sizing (name, vCPUs, hugepages, subsystems, stripe, drive format) and creates the request; the operator runs the discovery on the site through OCM. - The detail shows the draft as the site wrote it next to the requested sizing; "Size the draft" changes the sizing, "Approve and deploy" flips the one-way approval (typed confirmation). - The console's ClusterRole (chart and plain manifests) may manage the kind and read ManagedClusters; the mock serves two fixtures (Online, Drafted) and its personas carry the rule. Checked with a jsdom smoke run of the built bundle against the mock: the layer lists both fixtures, the drafted one prompts for approval, the deploy dialog opens, the detail renders the cluster and its nodes, no script error. dist/ rebuilt with the Dockerfile's pinned babel (7.29.7). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…om the hub A site's storage cluster is built from objects on the site's API server (the OperatorOps discovery, the ClusterDeploymentConfig draft it writes, the StorageCluster the approved draft expands into), which a hub managing the site through OCM does not reach. StorageSiteDeployment is the hub-side request: the site, the discovery filters, the sizing, and the one-way approval. Its controller (hub-only, registered with TestFailover when the ManifestWork API is served) carries the request through one ManifestWork in the site's hub namespace -- the discovery first, then a server-side apply of the sizing and the approval onto the draft once it names nodes -- and projects the draft, the StorageCluster and its nodes back through ManagedClusterViews. Phases Pending, Discovering, Drafted, Deploying, Online, Failed; conditions Delivered, Discovered, Approved, Ready; the cluster's uuid and pool are in the status, which is what a StorageClass names. Deleting the request orphans the work's resources: the storage cluster is never torn down by withdrawing the request. The console's ClusterRole may manage the kind and read ManagedClusters. The TestFailover CRD copies are the generator's rendering (as on #618). Design: simplyblock-dr docs/design/control-center-managed-discovery.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
schmidt-scaled
force-pushed
the
feat/storage-site-deployment
branch
from
October 3, 2026 12:55
bd9d80e to
8233d4b
Compare
…/install.yaml Lint (lll) on the chain-walk change and its tests; the installer bundle carries the TestFailover CRD's sourceVolumeMode. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… CRD Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A hub-side kind,
StorageSiteDeployment(storage.simplyblock.io/v1alpha2, namespaced), and its controller. One object per managed site requests that site's storage cluster from the hub: the discovery to run, the sizing of the draft it writes, and the approval that expands the draft.OperatorOpsdiscovery. Once the site's draft names nodes, the work also carries a server-side apply (field managerhub-deploy) of the sizing and the approval onto the draft. ManagedClusterViews project the draft, the StorageCluster and its StorageNodes back. The hub never holds a site kubeconfig, the same rule TestFailover follows.Why
Step a) of the DR UI run (discover a managed site's nodes, size the draft, approve it) could not be done from the hub's Control Center: those objects live on the site's API server. Design:
docs/design/control-center-managed-discovery.mdin simplyblock-dr.Base
This branch stacks on
realbed/control-center-on-csi-addons(ad4bad5), which is #618 at its tip (f8760ac) plus the Control Center chart templates of #512. The diff againstintegrate_csi_addonstherefore includes those commits; the change of this PR is the last commit, 8233d4b.Tests
storagesitedeployment_controller_unit_test.go, seven tests on the fake client:go build ./...andgo veton the controller, API and cmd packages pass.go test ./internal/controller/ ./internal/upgrade/... ./api/...passes, except the envtest suiteTestControllers, which needs etcd and was not runnable locally; CI runs it.🤖 Generated with Claude Code